Wikitech labswiki https://wikitech.wikimedia.org/wiki/Main_Page MediaWiki 1.47.0-wmf.14 first-letter Media Special Talk User User talk Wikitech Wikitech talk File File talk MediaWiki MediaWiki talk Template Template talk Help Help talk Category Category talk Obsolete Obsolete talk OfficeIT OfficeIT talk Tool Tool talk Nova Resource Nova Resource Talk Heira Heira Talk TimedText TimedText talk Module Module talk Deployments 0 4108 2445263 2445227 2026-08-09T21:58:15Z Superpes15 24252 +5 2445263 wikitext text/x-wiki {{Navigation MediaWiki deployment}} This page tracks '''upcoming''' '''deployments''' of software to the [[:m:Special:SiteMatrix|Wikimedia Foundation servers]]. == Getting started == Ensure you joined the {{irc|wikimedia-operations}} IRC channel as all deployment-related communications happen there. If you need help, contact [[:mw:Wikimedia Release Engineering Team|Release Engineering]] on IRC at {{irc|wikimedia-releng}}; and ping Tyler (<code>thcipriani</code>). * '''MediaWiki is deployed weekly''' through the [[/Train|Deployment Train]]. Other services follow their own schedule. * '''Times are pinned to San Francisco''', thus the UTC time changes in March and November per [[:en:Daylight saving time in the United States|DST]]. * '''Prefer regular [[Backport windows]]''' over adding new windows. To request deployment of a config change or backport, add your username and Gerrit URL to one of the backport windows on this page. You must be online in #wikimedia-operations on IRC during your deployment and install [[WikimediaDebug]] ahead of time. The #wikimedia-operations channel requires you to [[:m:IRC/Instructions#Register your nickname, identify, and enforce|register your nickname]] before you can join. ** You can use the '''backport scheduling tool''' to more easily edit this page: <div style="text-align: center; margin: 1em 0">{{Clickable button 2|:toollabs:schedule-deployment|Schedule a backport|class=mw-ui-progressive}}</div> * Tasks that meet [[/Inclusion criteria|Inclusion criteria]] '''require their own windows''', which includes long-running tasks. '''Schedule more time''' than you think you need to account for delays and set backs, we recommend one hour for most tasks. **To create or modify a recurring deploy window, send a patchset to [[:gitlab:repos/releng/release/-/blob/main/make-deployment-calendar/deployments-calendar.yaml|deployments-calendar.yaml file]] in <code>repos/releng/release.git</code>. **To create an one-off window, simply edit this page accordingly ** '''Announce''' changes to the [[mail:ops|ops mailing list]] ahead of time if you anticipate or are uncertain about noticeable impacts to database load, HTTP caching, or the introduction of new cookies. ** '''Announce''' deployments of major features to the community via [[:m:Tech/News/Next|Tech News]] and/or via other [[:mw:Wikimedia_Product_Guidance/Communication_channels|Product communication channels]]. * '''Something went wrong?''' See [[Incident response]]. Is there a user-impacting problem? Communicate in the {{irc|wikimedia-operations}} IRC channel. If there is a Phabricator task, ensure [[:phab:tag/wikimedia-incident/|#Wikimedia-Incident]] is tagged, and consider setting the [[:mw:Phabricator/Project_management#Priority_levels|Unbreak Now]] priority. __TOC__ {{anchor|Next Week|Near Term|Near term|Near-term}}{{clear}} [[Category:Deployment]] {{Note|content=Subscribe in Google Calendar via <code>wikimedia.org_rudis09ii2mm5fk4hgdjeh1u64@group.calendar.google.com</code>.<br>This may not include one-off windows. '''If there are differences, then the wiki page is canonical and correct'''.}} ==Week of August 10== ==={{Deployment_day|date=2026-08-09}}=== {{Deployment calendar event card |when=2026-08-09 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==={{Deployment_day|date=2026-08-10}}=== {{Deployment calendar event card |when=2026-08-10 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-10 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-10 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|Superpes|Superpes15}} {{deploy|type=config|gerrit=1323312|title=[frwiktionary] Add new Schème namespace and its talk|status=}} - {{phabricator|T415716}} {{deploy|type=config|gerrit=1323350|title=[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy)|status=}} - {{phabricator|T415307}} {{deploy|type=config|gerrit=1323779|title=[ukwiki] Remove reviewer usergroup|status=}} - {{phabricator|T434252}} {{deploy|type=config|gerrit=1322961|title=[slwiki] Reverting temporary logo for Wikipedia 25 (Vector legacy + Vector 2022)|status=}} - {{phabricator|T414265}} {{deploy|type=config|gerrit=1323827|title=[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy)|status=}} - {{phabricator|T414320}} {{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-10 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-10 08:30 SF |length=0.5 |window=Wikimedia Portals Update |who={{ircnick|jan_drewniak|Jan Drewniak}} |what=Weekly window for the portals page: https://www.wikipedia.org/ }} {{Deployment calendar event card |when=2026-08-10 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-10 10:00 SF |length=0.5 |window=Wikidata Query Service weekly deploy |who={{ircnick|ryankemper|Ryan}} |what=... }} {{Deployment calendar event card |when=2026-08-10 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-10 14:00 SF |length=2 |window=Weekly Security deployment window |who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}} |what=Held deployment window for Security-team related deploys. }} {{Deployment calendar event card |when=2026-08-10 16:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-10 19:00 SF |length=1 |window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Branch <code>wmf/1.47.0-wmf.15</code> }} {{Deployment calendar event card |when=2026-08-10 20:00 SF |length=1 |window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Deploy <code>wmf/1.47.0-wmf.15</code> to testwikis }} {{Deployment calendar event card |when=2026-08-10 21:00 SF |length=1 |window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) |who=N/A |what=Runs <code>scap clean auto</code> }} {{Deployment calendar event card |when=2026-08-10 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-10 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|Amir1|Amir}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-08-11}}=== {{Deployment calendar event card |when=2026-08-11 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-11 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-11 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-08-11 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-11 07:00 SF |length=0.5 |window=Test Kitchen UI Deployment Window |who=Experimentation Platform Team |what=Deployment of Test Kitchen UI (fka MPIC) }} {{Deployment calendar event card |when=2026-08-11 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-11 08:00 SF |length=1 |window=SRE Collaboration Services office hours |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=Services including Gerrit, Phorge (Phabricator), GitLab }} {{Deployment calendar event card |when=2026-08-11 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-08-11 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-11 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|brennen|Brennen}}, {{ircnick|jnuche|Jaime}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.14->1.47.0-wmf.15|1.47.0-wmf.14|1.47.0-wmf.14}} * group0 to [[mw:MediaWiki_1.47/wmf.15|1.47.0-wmf.15]] * '''Blockers: {{phabricator|T430834}}''' }} {{Deployment calendar event card |when=2026-08-11 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-11 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-11 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-08-12}}=== {{Deployment calendar event card |when=2026-08-12 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-12 01:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot) |who={{ircnick|brennen|Brennen}}, {{ircnick|jnuche|Jaime}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.15|1.47.0-wmf.14->1.47.0-wmf.15|1.47.0-wmf.14}} * group1 to [[mw:MediaWiki_1.47/wmf.15|1.47.0-wmf.15]] * '''Blockers: {{phabricator|T430834}}''' }} {{Deployment calendar event card |when=2026-08-12 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-12 04:00 SF |length=1 |window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]] |who=Marielle ({{ircnick|mvolz}}) |what=See [[mw:Citoid|Citoid]] }} {{Deployment calendar event card |when=2026-08-12 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-12 07:00 SF |length=1 |window=Wikifunctions Services UTC Afternoon |who=Abstract Wikipedia team (Africa, Europe, Eastern Americas) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-08-12 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-12 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-12 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|brennen|Brennen}}, {{ircnick|jnuche|Jaime}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.15|1.47.0-wmf.14->1.47.0-wmf.15|1.47.0-wmf.14}} * group1 to [[mw:MediaWiki_1.47/wmf.15|1.47.0-wmf.15]] * '''Blockers: {{phabricator|T430834}}''' }} {{Deployment calendar event card |when=2026-08-12 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-12 14:00 SF |length=1 |window=Wikifunctions Services UTC Late |who=Abstract Wikipedia team (North and South America) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-08-12 15:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-12 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-12 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|Amir1|Amir}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-08-13}}=== {{Deployment calendar event card |when=2026-08-13 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-13 01:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot) |who={{ircnick|brennen|Brennen}}, {{ircnick|jnuche|Jaime}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.15|1.47.0-wmf.15|1.47.0-wmf.14->1.47.0-wmf.15}} * group2 to [[mw:MediaWiki_1.47/wmf.15|1.47.0-wmf.15]] * '''Blockers: {{phabricator|T430834}}''' }} {{Deployment calendar event card |when=2026-08-13 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-13 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-08-13 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-13 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-13 08:00 SF |length=1 |window=Train log triage |who={{ircnick|brennen|Brennen}}, {{ircnick|jnuche|Jaime}} |what=See [[Heterogeneous deployment/Train deploys#Breakage]] }} {{Deployment calendar event card |when=2026-08-13 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-08-13 10:00 SF |length=1 |window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) |who={{ircnick|bd808}} |what=... }} {{Deployment calendar event card |when=2026-08-13 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-13 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|brennen|Brennen}}, {{ircnick|jnuche|Jaime}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.15|1.47.0-wmf.15|1.47.0-wmf.14->1.47.0-wmf.15}} * group2 to [[mw:MediaWiki_1.47/wmf.15|1.47.0-wmf.15]] * '''Blockers: {{phabricator|T430834}}''' }} {{Deployment calendar event card |when=2026-08-13 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-13 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-13 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-08-14}}=== {{Deployment calendar event card |when=2026-08-14 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} {{Deployment calendar event card |when=2026-08-14 04:00 SF |length=0.5 |window=GitLab version upgrades |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=GitLab version upgrades }} ==={{Deployment_day|date=2026-08-15}}=== {{Deployment calendar event card |when=2026-08-15 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==Week of August 17== ==={{Deployment_day|date=2026-08-16}}=== {{Deployment calendar event card |when=2026-08-16 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==={{Deployment_day|date=2026-08-17}}=== {{Deployment calendar event card |when=2026-08-17 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-17 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-17 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-17 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-17 08:30 SF |length=0.5 |window=Wikimedia Portals Update |who={{ircnick|jan_drewniak|Jan Drewniak}} |what=Weekly window for the portals page: https://www.wikipedia.org/ }} {{Deployment calendar event card |when=2026-08-17 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-17 10:00 SF |length=0.5 |window=Wikidata Query Service weekly deploy |who={{ircnick|ryankemper|Ryan}} |what=... }} {{Deployment calendar event card |when=2026-08-17 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-17 14:00 SF |length=2 |window=Weekly Security deployment window |who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}} |what=Held deployment window for Security-team related deploys. }} {{Deployment calendar event card |when=2026-08-17 16:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-17 19:00 SF |length=1 |window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Branch <code>wmf/1.47.0-wmf.16</code> }} {{Deployment calendar event card |when=2026-08-17 20:00 SF |length=1 |window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Deploy <code>wmf/1.47.0-wmf.16</code> to testwikis }} {{Deployment calendar event card |when=2026-08-17 21:00 SF |length=1 |window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) |who=N/A |what=Runs <code>scap clean auto</code> }} {{Deployment calendar event card |when=2026-08-17 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-17 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|Amir1|Amir}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-08-18}}=== {{Deployment calendar event card |when=2026-08-18 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-18 01:00 SF |length=2 |window=MediaWiki train - Utc-0+Utc-7 Version |who={{ircnick|andre|Andre}}, {{ircnick|brennen|Brennen}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.15->1.47.0-wmf.16|1.47.0-wmf.15|1.47.0-wmf.15}} * group0 to [[mw:MediaWiki_1.47/wmf.16|1.47.0-wmf.16]] * '''Blockers: {{phabricator|T430835}}''' }} {{Deployment calendar event card |when=2026-08-18 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-18 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-08-18 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-18 07:00 SF |length=0.5 |window=Test Kitchen UI Deployment Window |who=Experimentation Platform Team |what=Deployment of Test Kitchen UI (fka MPIC) }} {{Deployment calendar event card |when=2026-08-18 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-18 08:00 SF |length=1 |window=SRE Collaboration Services office hours |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=Services including Gerrit, Phorge (Phabricator), GitLab }} {{Deployment calendar event card |when=2026-08-18 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-08-18 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-18 11:00 SF |length=2 |window=MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) |who={{ircnick|andre|Andre}}, {{ircnick|brennen|Brennen}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.15->1.47.0-wmf.16|1.47.0-wmf.15|1.47.0-wmf.15}} * group0 to [[mw:MediaWiki_1.47/wmf.16|1.47.0-wmf.16]] * '''Blockers: {{phabricator|T430835}}''' }} {{Deployment calendar event card |when=2026-08-18 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-18 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-18 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-08-19}}=== {{Deployment calendar event card |when=2026-08-19 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-19 01:00 SF |length=2 |window=MediaWiki train - Utc-0+Utc-7 Version |who={{ircnick|andre|Andre}}, {{ircnick|brennen|Brennen}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.16|1.47.0-wmf.15->1.47.0-wmf.16|1.47.0-wmf.15}} * group1 to [[mw:MediaWiki_1.47/wmf.16|1.47.0-wmf.16]] * '''Blockers: {{phabricator|T430835}}''' }} {{Deployment calendar event card |when=2026-08-19 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-19 04:00 SF |length=1 |window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]] |who=Marielle ({{ircnick|mvolz}}) |what=See [[mw:Citoid|Citoid]] }} {{Deployment calendar event card |when=2026-08-19 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-19 07:00 SF |length=1 |window=Wikifunctions Services UTC Afternoon |who=Abstract Wikipedia team (Africa, Europe, Eastern Americas) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-08-19 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-19 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-19 11:00 SF |length=2 |window=MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) |who={{ircnick|andre|Andre}}, {{ircnick|brennen|Brennen}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.16|1.47.0-wmf.15->1.47.0-wmf.16|1.47.0-wmf.15}} * group1 to [[mw:MediaWiki_1.47/wmf.16|1.47.0-wmf.16]] * '''Blockers: {{phabricator|T430835}}''' }} {{Deployment calendar event card |when=2026-08-19 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-19 14:00 SF |length=1 |window=Wikifunctions Services UTC Late |who=Abstract Wikipedia team (North and South America) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-08-19 15:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-19 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-19 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|Amir1|Amir}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-08-20}}=== {{Deployment calendar event card |when=2026-08-20 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-20 01:00 SF |length=2 |window=MediaWiki train - Utc-0+Utc-7 Version |who={{ircnick|andre|Andre}}, {{ircnick|brennen|Brennen}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.16|1.47.0-wmf.16|1.47.0-wmf.15->1.47.0-wmf.16}} * group2 to [[mw:MediaWiki_1.47/wmf.16|1.47.0-wmf.16]] * '''Blockers: {{phabricator|T430835}}''' }} {{Deployment calendar event card |when=2026-08-20 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-20 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-08-20 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-20 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-20 08:00 SF |length=1 |window=Train log triage |who={{ircnick|andre|Andre}}, {{ircnick|brennen|Brennen}} |what=See [[Heterogeneous deployment/Train deploys#Breakage]] }} {{Deployment calendar event card |when=2026-08-20 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-08-20 10:00 SF |length=1 |window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) |who={{ircnick|bd808}} |what=... }} {{Deployment calendar event card |when=2026-08-20 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-20 11:00 SF |length=2 |window=MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) |who={{ircnick|andre|Andre}}, {{ircnick|brennen|Brennen}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.16|1.47.0-wmf.16|1.47.0-wmf.15->1.47.0-wmf.16}} * group2 to [[mw:MediaWiki_1.47/wmf.16|1.47.0-wmf.16]] * '''Blockers: {{phabricator|T430835}}''' }} {{Deployment calendar event card |when=2026-08-20 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-20 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-20 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-08-21}}=== {{Deployment calendar event card |when=2026-08-21 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} {{Deployment calendar event card |when=2026-08-21 04:00 SF |length=0.5 |window=GitLab version upgrades |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=GitLab version upgrades }} ==={{Deployment_day|date=2026-08-22}}=== {{Deployment calendar event card |when=2026-08-22 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} 0rqspkkprfv85cavi494gjrdqomay42 2445274 2445263 2026-08-10T10:13:55Z ScheduleDeploymentBot 37566 Add [[gerrit:1323939]] to Monday, August 10 UTC afternoon backport window 2445274 wikitext text/x-wiki {{Navigation MediaWiki deployment}} This page tracks '''upcoming''' '''deployments''' of software to the [[:m:Special:SiteMatrix|Wikimedia Foundation servers]]. == Getting started == Ensure you joined the {{irc|wikimedia-operations}} IRC channel as all deployment-related communications happen there. If you need help, contact [[:mw:Wikimedia Release Engineering Team|Release Engineering]] on IRC at {{irc|wikimedia-releng}}; and ping Tyler (<code>thcipriani</code>). * '''MediaWiki is deployed weekly''' through the [[/Train|Deployment Train]]. Other services follow their own schedule. * '''Times are pinned to San Francisco''', thus the UTC time changes in March and November per [[:en:Daylight saving time in the United States|DST]]. * '''Prefer regular [[Backport windows]]''' over adding new windows. To request deployment of a config change or backport, add your username and Gerrit URL to one of the backport windows on this page. You must be online in #wikimedia-operations on IRC during your deployment and install [[WikimediaDebug]] ahead of time. The #wikimedia-operations channel requires you to [[:m:IRC/Instructions#Register your nickname, identify, and enforce|register your nickname]] before you can join. ** You can use the '''backport scheduling tool''' to more easily edit this page: <div style="text-align: center; margin: 1em 0">{{Clickable button 2|:toollabs:schedule-deployment|Schedule a backport|class=mw-ui-progressive}}</div> * Tasks that meet [[/Inclusion criteria|Inclusion criteria]] '''require their own windows''', which includes long-running tasks. '''Schedule more time''' than you think you need to account for delays and set backs, we recommend one hour for most tasks. **To create or modify a recurring deploy window, send a patchset to [[:gitlab:repos/releng/release/-/blob/main/make-deployment-calendar/deployments-calendar.yaml|deployments-calendar.yaml file]] in <code>repos/releng/release.git</code>. **To create an one-off window, simply edit this page accordingly ** '''Announce''' changes to the [[mail:ops|ops mailing list]] ahead of time if you anticipate or are uncertain about noticeable impacts to database load, HTTP caching, or the introduction of new cookies. ** '''Announce''' deployments of major features to the community via [[:m:Tech/News/Next|Tech News]] and/or via other [[:mw:Wikimedia_Product_Guidance/Communication_channels|Product communication channels]]. * '''Something went wrong?''' See [[Incident response]]. Is there a user-impacting problem? Communicate in the {{irc|wikimedia-operations}} IRC channel. If there is a Phabricator task, ensure [[:phab:tag/wikimedia-incident/|#Wikimedia-Incident]] is tagged, and consider setting the [[:mw:Phabricator/Project_management#Priority_levels|Unbreak Now]] priority. __TOC__ {{anchor|Next Week|Near Term|Near term|Near-term}}{{clear}} [[Category:Deployment]] {{Note|content=Subscribe in Google Calendar via <code>wikimedia.org_rudis09ii2mm5fk4hgdjeh1u64@group.calendar.google.com</code>.<br>This may not include one-off windows. '''If there are differences, then the wiki page is canonical and correct'''.}} ==Week of August 10== ==={{Deployment_day|date=2026-08-09}}=== {{Deployment calendar event card |when=2026-08-09 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==={{Deployment_day|date=2026-08-10}}=== {{Deployment calendar event card |when=2026-08-10 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-10 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-10 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|Superpes|Superpes15}} {{deploy|type=config|gerrit=1323312|title=[frwiktionary] Add new Schème namespace and its talk|status=}} - {{phabricator|T415716}} {{deploy|type=config|gerrit=1323350|title=[tgwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy)|status=}} - {{phabricator|T415307}} {{deploy|type=config|gerrit=1323779|title=[ukwiki] Remove reviewer usergroup|status=}} - {{phabricator|T434252}} {{deploy|type=config|gerrit=1322961|title=[slwiki] Reverting temporary logo for Wikipedia 25 (Vector legacy + Vector 2022)|status=}} - {{phabricator|T414265}} {{deploy|type=config|gerrit=1323827|title=[itwiki] Remove temporary logo for Wikipedia 25 (Vector 2022 + Vector Legacy)|status=}} - {{phabricator|T414320}} {{ircnick|WMDE-Fisch|WMDE-Fisch}} {{deploy|type=config|gerrit=1323939|title=Enable sub-references on more group2 wikis (batch3)|status=}} {{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-10 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-10 08:30 SF |length=0.5 |window=Wikimedia Portals Update |who={{ircnick|jan_drewniak|Jan Drewniak}} |what=Weekly window for the portals page: https://www.wikipedia.org/ }} {{Deployment calendar event card |when=2026-08-10 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-10 10:00 SF |length=0.5 |window=Wikidata Query Service weekly deploy |who={{ircnick|ryankemper|Ryan}} |what=... }} {{Deployment calendar event card |when=2026-08-10 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-10 14:00 SF |length=2 |window=Weekly Security deployment window |who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}} |what=Held deployment window for Security-team related deploys. }} {{Deployment calendar event card |when=2026-08-10 16:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-10 19:00 SF |length=1 |window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Branch <code>wmf/1.47.0-wmf.15</code> }} {{Deployment calendar event card |when=2026-08-10 20:00 SF |length=1 |window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Deploy <code>wmf/1.47.0-wmf.15</code> to testwikis }} {{Deployment calendar event card |when=2026-08-10 21:00 SF |length=1 |window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) |who=N/A |what=Runs <code>scap clean auto</code> }} {{Deployment calendar event card |when=2026-08-10 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-10 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|Amir1|Amir}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-08-11}}=== {{Deployment calendar event card |when=2026-08-11 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-11 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-11 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-08-11 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-11 07:00 SF |length=0.5 |window=Test Kitchen UI Deployment Window |who=Experimentation Platform Team |what=Deployment of Test Kitchen UI (fka MPIC) }} {{Deployment calendar event card |when=2026-08-11 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-11 08:00 SF |length=1 |window=SRE Collaboration Services office hours |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=Services including Gerrit, Phorge (Phabricator), GitLab }} {{Deployment calendar event card |when=2026-08-11 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-08-11 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-11 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|brennen|Brennen}}, {{ircnick|jnuche|Jaime}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.14->1.47.0-wmf.15|1.47.0-wmf.14|1.47.0-wmf.14}} * group0 to [[mw:MediaWiki_1.47/wmf.15|1.47.0-wmf.15]] * '''Blockers: {{phabricator|T430834}}''' }} {{Deployment calendar event card |when=2026-08-11 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-11 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-11 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-08-12}}=== {{Deployment calendar event card |when=2026-08-12 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-12 01:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot) |who={{ircnick|brennen|Brennen}}, {{ircnick|jnuche|Jaime}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.15|1.47.0-wmf.14->1.47.0-wmf.15|1.47.0-wmf.14}} * group1 to [[mw:MediaWiki_1.47/wmf.15|1.47.0-wmf.15]] * '''Blockers: {{phabricator|T430834}}''' }} {{Deployment calendar event card |when=2026-08-12 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-12 04:00 SF |length=1 |window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]] |who=Marielle ({{ircnick|mvolz}}) |what=See [[mw:Citoid|Citoid]] }} {{Deployment calendar event card |when=2026-08-12 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-12 07:00 SF |length=1 |window=Wikifunctions Services UTC Afternoon |who=Abstract Wikipedia team (Africa, Europe, Eastern Americas) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-08-12 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-12 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-12 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|brennen|Brennen}}, {{ircnick|jnuche|Jaime}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.15|1.47.0-wmf.14->1.47.0-wmf.15|1.47.0-wmf.14}} * group1 to [[mw:MediaWiki_1.47/wmf.15|1.47.0-wmf.15]] * '''Blockers: {{phabricator|T430834}}''' }} {{Deployment calendar event card |when=2026-08-12 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-12 14:00 SF |length=1 |window=Wikifunctions Services UTC Late |who=Abstract Wikipedia team (North and South America) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-08-12 15:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-12 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-12 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|Amir1|Amir}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-08-13}}=== {{Deployment calendar event card |when=2026-08-13 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-13 01:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot) |who={{ircnick|brennen|Brennen}}, {{ircnick|jnuche|Jaime}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.15|1.47.0-wmf.15|1.47.0-wmf.14->1.47.0-wmf.15}} * group2 to [[mw:MediaWiki_1.47/wmf.15|1.47.0-wmf.15]] * '''Blockers: {{phabricator|T430834}}''' }} {{Deployment calendar event card |when=2026-08-13 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-13 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-08-13 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-13 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-13 08:00 SF |length=1 |window=Train log triage |who={{ircnick|brennen|Brennen}}, {{ircnick|jnuche|Jaime}} |what=See [[Heterogeneous deployment/Train deploys#Breakage]] }} {{Deployment calendar event card |when=2026-08-13 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-08-13 10:00 SF |length=1 |window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) |who={{ircnick|bd808}} |what=... }} {{Deployment calendar event card |when=2026-08-13 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-13 11:00 SF |length=2 |window=MediaWiki train - Utc-7+Utc-0 Version |who={{ircnick|brennen|Brennen}}, {{ircnick|jnuche|Jaime}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.15|1.47.0-wmf.15|1.47.0-wmf.14->1.47.0-wmf.15}} * group2 to [[mw:MediaWiki_1.47/wmf.15|1.47.0-wmf.15]] * '''Blockers: {{phabricator|T430834}}''' }} {{Deployment calendar event card |when=2026-08-13 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-13 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-13 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-08-14}}=== {{Deployment calendar event card |when=2026-08-14 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} {{Deployment calendar event card |when=2026-08-14 04:00 SF |length=0.5 |window=GitLab version upgrades |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=GitLab version upgrades }} ==={{Deployment_day|date=2026-08-15}}=== {{Deployment calendar event card |when=2026-08-15 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==Week of August 17== ==={{Deployment_day|date=2026-08-16}}=== {{Deployment calendar event card |when=2026-08-16 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} ==={{Deployment_day|date=2026-08-17}}=== {{Deployment calendar event card |when=2026-08-17 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-17 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-17 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-17 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-17 08:30 SF |length=0.5 |window=Wikimedia Portals Update |who={{ircnick|jan_drewniak|Jan Drewniak}} |what=Weekly window for the portals page: https://www.wikipedia.org/ }} {{Deployment calendar event card |when=2026-08-17 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-17 10:00 SF |length=0.5 |window=Wikidata Query Service weekly deploy |who={{ircnick|ryankemper|Ryan}} |what=... }} {{Deployment calendar event card |when=2026-08-17 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-17 14:00 SF |length=2 |window=Weekly Security deployment window |who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}} |what=Held deployment window for Security-team related deploys. }} {{Deployment calendar event card |when=2026-08-17 16:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-17 19:00 SF |length=1 |window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Branch <code>wmf/1.47.0-wmf.16</code> }} {{Deployment calendar event card |when=2026-08-17 20:00 SF |length=1 |window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]] |who=N/A |what=Deploy <code>wmf/1.47.0-wmf.16</code> to testwikis }} {{Deployment calendar event card |when=2026-08-17 21:00 SF |length=1 |window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) |who=N/A |what=Runs <code>scap clean auto</code> }} {{Deployment calendar event card |when=2026-08-17 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-17 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|Amir1|Amir}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-08-18}}=== {{Deployment calendar event card |when=2026-08-18 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-18 01:00 SF |length=2 |window=MediaWiki train - Utc-0+Utc-7 Version |who={{ircnick|andre|Andre}}, {{ircnick|brennen|Brennen}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.15->1.47.0-wmf.16|1.47.0-wmf.15|1.47.0-wmf.15}} * group0 to [[mw:MediaWiki_1.47/wmf.16|1.47.0-wmf.16]] * '''Blockers: {{phabricator|T430835}}''' }} {{Deployment calendar event card |when=2026-08-18 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-18 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-08-18 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-18 07:00 SF |length=0.5 |window=Test Kitchen UI Deployment Window |who=Experimentation Platform Team |what=Deployment of Test Kitchen UI (fka MPIC) }} {{Deployment calendar event card |when=2026-08-18 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-18 08:00 SF |length=1 |window=SRE Collaboration Services office hours |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=Services including Gerrit, Phorge (Phabricator), GitLab }} {{Deployment calendar event card |when=2026-08-18 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-08-18 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-18 11:00 SF |length=2 |window=MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) |who={{ircnick|andre|Andre}}, {{ircnick|brennen|Brennen}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.15->1.47.0-wmf.16|1.47.0-wmf.15|1.47.0-wmf.15}} * group0 to [[mw:MediaWiki_1.47/wmf.16|1.47.0-wmf.16]] * '''Blockers: {{phabricator|T430835}}''' }} {{Deployment calendar event card |when=2026-08-18 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-18 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-18 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-08-19}}=== {{Deployment calendar event card |when=2026-08-19 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-19 01:00 SF |length=2 |window=MediaWiki train - Utc-0+Utc-7 Version |who={{ircnick|andre|Andre}}, {{ircnick|brennen|Brennen}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.16|1.47.0-wmf.15->1.47.0-wmf.16|1.47.0-wmf.15}} * group1 to [[mw:MediaWiki_1.47/wmf.16|1.47.0-wmf.16]] * '''Blockers: {{phabricator|T430835}}''' }} {{Deployment calendar event card |when=2026-08-19 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-19 04:00 SF |length=1 |window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]] |who=Marielle ({{ircnick|mvolz}}) |what=See [[mw:Citoid|Citoid]] }} {{Deployment calendar event card |when=2026-08-19 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-19 07:00 SF |length=1 |window=Wikifunctions Services UTC Afternoon |who=Abstract Wikipedia team (Africa, Europe, Eastern Americas) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-08-19 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-19 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-19 11:00 SF |length=2 |window=MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) |who={{ircnick|andre|Andre}}, {{ircnick|brennen|Brennen}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.16|1.47.0-wmf.15->1.47.0-wmf.16|1.47.0-wmf.15}} * group1 to [[mw:MediaWiki_1.47/wmf.16|1.47.0-wmf.16]] * '''Blockers: {{phabricator|T430835}}''' }} {{Deployment calendar event card |when=2026-08-19 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-19 14:00 SF |length=1 |window=Wikifunctions Services UTC Late |who=Abstract Wikipedia team (North and South America) |what=Wikifunctions back-end k8s services }} {{Deployment calendar event card |when=2026-08-19 15:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-19 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-19 23:00 SF |length=0.5 |window=Primary database switchover |who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|Amir1|Amir}}, {{ircnick|federico3|Federico Ceratto}} |what=Held deployment window for database primary masters maintenance }} ==={{Deployment_day|date=2026-08-20}}=== {{Deployment calendar event card |when=2026-08-20 00:00 SF |length=1 |window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-20 01:00 SF |length=2 |window=MediaWiki train - Utc-0+Utc-7 Version |who={{ircnick|andre|Andre}}, {{ircnick|brennen|Brennen}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.16|1.47.0-wmf.16|1.47.0-wmf.15->1.47.0-wmf.16}} * group2 to [[mw:MediaWiki_1.47/wmf.16|1.47.0-wmf.16]] * '''Blockers: {{phabricator|T430835}}''' }} {{Deployment calendar event card |when=2026-08-20 03:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-20 05:00 SF |length=1 |window=Mobileapps/RESTBase/Wikifeeds |who=Content Transform Team |what=Content transform team node services (mobileapps/wikifeeds) }} {{Deployment calendar event card |when=2026-08-20 06:00 SF |length=1 |window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-20 07:30 SF |length=0.5 |window=Test Kitchen Experiment Deployment Window |who=Test Kitchen |what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]]. }} {{Deployment calendar event card |when=2026-08-20 08:00 SF |length=1 |window=Train log triage |who={{ircnick|andre|Andre}}, {{ircnick|brennen|Brennen}} |what=See [[Heterogeneous deployment/Train deploys#Breakage]] }} {{Deployment calendar event card |when=2026-08-20 09:00 SF |length=1 |window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small> |who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to Puppet change'' }} {{Deployment calendar event card |when=2026-08-20 10:00 SF |length=1 |window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) |who={{ircnick|bd808}} |what=... }} {{Deployment calendar event card |when=2026-08-20 10:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} {{Deployment calendar event card |when=2026-08-20 11:00 SF |length=2 |window=MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) |who={{ircnick|andre|Andre}}, {{ircnick|brennen|Brennen}} |what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]] {{DeployOneWeekMini|1.47.0-wmf.16|1.47.0-wmf.16|1.47.0-wmf.15->1.47.0-wmf.16}} * group2 to [[mw:MediaWiki_1.47/wmf.16|1.47.0-wmf.16]] * '''Blockers: {{phabricator|T430835}}''' }} {{Deployment calendar event card |when=2026-08-20 13:00 SF |length=1 |window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small> |who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}} |what={{ircnick|irc-nickname|Requesting Developer}} * ''Gerrit link to backport or config change'' }} {{Deployment calendar event card |when=2026-08-20 14:00 SF |length=1 |window=Readers deployment window |who=Readers |what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start }} {{Deployment calendar event card |when=2026-08-20 23:00 SF |length=1 |window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early) |who=SRE team |what=MediaWiki-related infrastructure changes that need a kubernetes deployment. }} ==={{Deployment_day|date=2026-08-21}}=== {{Deployment calendar event card |when=2026-08-21 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} {{Deployment calendar event card |when=2026-08-21 04:00 SF |length=0.5 |window=GitLab version upgrades |who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}} |what=GitLab version upgrades }} ==={{Deployment_day|date=2026-08-22}}=== {{Deployment calendar event card |when=2026-08-22 00:00 SF |length=24 |window=No deploys all day! See [[Deployments/Emergencies]] if things are broken. |who= |what=No Deploys }} t1sz0f77g596jh64ljsrna15ertx4pc Server Admin Log 0 7919 2445256 2445247 2026-08-09T16:01:00Z Stashbot 7414 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply 2445256 wikitext text/x-wiki == 2026-08-09 == * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> tt141o4jjplixarh5nryg90e2z1acbw 2445257 2445256 2026-08-09T16:01:03Z Stashbot 7414 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply 2445257 wikitext text/x-wiki == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> 5d7obtph4efsp1t6gkusrcitjtydfkh 2445258 2445257 2026-08-09T16:01:04Z Stashbot 7414 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply 2445258 wikitext text/x-wiki == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> r5muyws88h3e2z11lmswbpfv140fs2n 2445259 2445258 2026-08-09T16:01:07Z Stashbot 7414 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply 2445259 wikitext text/x-wiki == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> 18ssxbdk4ucmthf5xjy22fdvt3uuiut 2445266 2445259 2026-08-10T02:00:16Z Stashbot 7414 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image 2445266 wikitext text/x-wiki == 2026-08-10 == * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> 3ppk6i45amgbyb8fohnjcvqy02u9k2x 2445267 2445266 2026-08-10T02:07:04Z Stashbot 7414 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) 2445267 wikitext text/x-wiki == 2026-08-10 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> etxra37z1pd481dei3qb1ak2f6ensqe 2445271 2445267 2026-08-10T07:35:28Z Stashbot 7414 _joe_: restarting squid on urldownloader1006 2445271 wikitext text/x-wiki == 2026-08-10 == * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> roxub4spy2aomtxah5pvz9c1hzwtlou 2445272 2445271 2026-08-10T07:57:12Z Stashbot 7414 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies 2445272 wikitext text/x-wiki == 2026-08-10 == * 07:57 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> 0e6j3i9beg83f5zygx9dlrlsy02xvr2 2445273 2445272 2026-08-10T07:57:26Z Stashbot 7414 hashar@deploy1003: Finished deploy [integration/docroot@7772132]: update build dependencies (duration: 00m 13s) 2445273 wikitext text/x-wiki == 2026-08-10 == * 07:57 hashar@deploy1003: Finished deploy [integration/docroot@7772132]: update build dependencies (duration: 00m 13s) * 07:57 hashar@deploy1003: Started deploy [integration/docroot@7772132]: update build dependencies * 07:35 _joe_: restarting squid on urldownloader1006 * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 48s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-09 == * 16:01 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:01 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:00 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-08 == * 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s) * 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s) * 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) * 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s) * 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s) * 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-07 == * 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie * 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]] * 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]] * 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]] * 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie * 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning * 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning * 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]] * 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. * 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'. * 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning * 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]] * 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]] * 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json * 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning * 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning * 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003 * 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet * 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]] * 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]] * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" * 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox * 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet * 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]] * 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission * 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-06 == * 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]] * 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]] * 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) * 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment * 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be * 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] * 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) * 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment * 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] * 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) * 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment * 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] * 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie * 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s) * 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 18:18 ladsgroup@deploy1003: Stopping before sync operations * 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) * 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie * 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage * 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage * 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]]) * 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]]) * 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]]) * 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage * 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage * 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm * 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm * 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]] * 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm * 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm * 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]] * 13:58 swfrench@dns1004: END - running authdns-update * 13:56 swfrench@dns1004: START - running authdns-update * 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm * 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . * 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage * 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply * 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update * 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update * 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json * 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json * 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command * 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning * 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning * 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json * 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet * 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]] * 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . * 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . * 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . * 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . * 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . * 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]] * 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning * 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-05 == * 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm * 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage * 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm * 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet * 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet * 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm * 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage * 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm * 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet * 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet * 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm * 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox * 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas * 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage * 20:24 cjming: end of UTC late backport window * 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) * 20:18 cjming@deploy1003: cjming: Continuing with deployment * 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] * 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) * 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm * 20:08 jforrester@deploy1003: jforrester: Continuing with deployment * 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] * 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]] * 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 * 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie * 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]] * 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm * 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet * 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage * 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet * 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie * 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . * 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" * 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox * 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad * 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]] * 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) * 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet * 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet * 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad * 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]]) * 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]]) * 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]]) * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet * 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet * 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) * 15:58 reedy@deploy1003: reedy: Continuing with deployment * 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] * 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 * 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 * 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]] * 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]] * 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" * 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]] * 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 * 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie * 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet * 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet * 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet * 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm * 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]] * 14:39 swfrench@dns1004: END - running authdns-update * 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:37 swfrench@dns1004: START - running authdns-update * 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie * 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie * 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage * 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad * 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad * 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage * 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage * 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm * 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie * 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 * 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 * 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" * 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) * 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox * 13:35 reedy@deploy1003: reedy: Continuing with deployment * 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] * 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 * 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie * 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet * 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet * 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet * 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet * 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie * 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage * 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 12:24 topranks: update bgp confed settings in eqsin * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 * 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 * 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" * 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie * 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox * 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . * 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 * 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie * 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet * 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet * 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046 * 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie * 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet * 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet * 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie * 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json * 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning * 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning * 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 * 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 * 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]] * 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" * 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning * 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox * 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning * 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache * 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning * 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 * 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json * 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json * 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet * 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet * 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox * 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json * 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet * 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]] * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 * 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 * 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 * 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]] * 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]] * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]] * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json * 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]] * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 * 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]] * 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]] * 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 * 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json * 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance * 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json * 07:21 slyngshede@dns1004: END - running authdns-update * 07:19 slyngshede@dns1004: START - running authdns-update * 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json * 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json * 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json * 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json * 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance * 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json * 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json * 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json * 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json * 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json * 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance * 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance * 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json * 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json * 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json * 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json * 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json * 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance * 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json * 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json * 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json * 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json * 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance * 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json * 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json * 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json * 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json == 2026-08-04 == * 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json * 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json * 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json * 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json * 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json * 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json * 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance * 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) * 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment * 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm * 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] * 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage * 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance * 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json * 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" * 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json * 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046 * 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm * 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json * 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json * 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}} * 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json * 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json * 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json * 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet * 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet * 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s) * 17:48 swfrench@deploy1003: swfrench: Continuing with deployment * 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] * 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json * 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json * 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) * 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure * 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image * 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) * 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue * 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet * 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit * 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet * 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ * 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 * 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json * 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 16:25 mutante: restarting gerrit - dropped outdated RSA host key * 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 * 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json * 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json * 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia * 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia * 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] * 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]] * 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie * 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s) * 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] * 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s) * 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]] * 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync * 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 * 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 * 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) * 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" * 14:30 otto@deploy1003: otto: Continuing with deployment * 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox * 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 * 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie * 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:13 swfrench@dns1004: END - running authdns-update * 14:13 Msz2001: Finished deployments for UTC afternoon backport window * 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) * 14:11 swfrench@dns1004: START - running authdns-update * 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment * 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser * 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] * 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]] * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet * 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet * 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) * 13:45 swfrench@dns1004: END - running authdns-update * 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment * 13:43 swfrench@dns1004: START - running authdns-update * 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] * 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet * 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet * 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply * 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 13:05 swfrench@dns1004: END - running authdns-update * 13:03 swfrench@dns1004: START - running authdns-update * 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom * 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning * 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie * 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie * 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage * 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet * 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet * 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet * 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie * 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet * 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet * 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie * 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage * 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet * 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet * 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie * 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet * 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows * 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet * 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 * 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" * 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet * 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet * 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad * 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet * 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet * 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage * 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 * 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie * 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet * 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet * 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 * 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet * 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet * 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie * 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet * 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet * 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet * 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet * 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet * 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet * 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet * 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie * 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie * 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie * 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie * 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet * 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie * 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage * 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie * 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage * 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage * 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage * 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie * 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage * 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 * 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 * 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie * 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 * 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 * 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 * 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 * 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie * 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" * 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie * 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie * 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 * 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie * 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet * 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie * 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet * 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet * 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie * 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage * 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage * 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071 * 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070 * 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069 * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org * 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org * 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet * 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie * 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet * 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie * 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet * 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet * 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie * 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet * 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet * 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet * 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage * 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage * 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply * 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet * 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie * 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet * 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet * 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet * 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie * 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet * 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet * 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie * 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet * 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet * 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie * 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet * 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]] * 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet * 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage * 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage * 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]] * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie * 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie * 06:50 slyngshede@dns1004: END - running authdns-update * 06:48 slyngshede@dns1004: START - running authdns-update * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) * 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s) * 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) * 00:41 cjming@deploy1003: cjming: Continuing with deployment * 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] == 2026-08-03 == * 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply * 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply * 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet * 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet * 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync * 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync * 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie * 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage * 21:42 dancy@deploy1003: Stopping before sync operations * 21:41 dancy@deploy1003: Started scap sync-world: testing * 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) * 21:33 dancy@deploy1003: dancy: Continuing with deployment * 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] * 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie * 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) * 20:59 dancy@deploy1003: dancy: Continuing with deployment * 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] * 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) * 20:48 cjming@deploy1003: cjming: Continuing with deployment * 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] * 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) * 20:38 arlolra@deploy1003: arlolra: Continuing with deployment * 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] * 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) * 20:12 krinkle@deploy1003: krinkle: Continuing with deployment * 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] * 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw * 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) * 18:53 krinkle@deploy1003: krinkle: Continuing with deployment * 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw * 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] * 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet * 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie * 18:33 krinkle@deploy1003: krinkle: Continuing with deployment * 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] * 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage * 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie * 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" * 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet * 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie * 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" * 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie * 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie * 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox * 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet * 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 * 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046 * 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage * 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage * 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie * 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie * 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie * 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage * 15:33 jhathaway@dns1004: END - running authdns-update * 15:31 jhathaway@dns1004: START - running authdns-update * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie * 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie * 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json * 15:09 dancy@deploy1003: Started scap sync-world: testing * 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts * 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s) * 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage * 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie * 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie * 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie * 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie * 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie * 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) * 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage * 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] * 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie * 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie * 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet * 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet * 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet * 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet * 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet * 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie * 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie * 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet * 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet * 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s) * 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage * 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage * 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update * 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie * 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie * 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet * 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet * 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie * 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network * 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:19 marostegui@dns1004: END - running authdns-update * 11:17 marostegui@dns1004: START - running authdns-update * 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply * 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply * 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply * 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply * 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply * 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply * 11:02 marostegui@dns1004: END - running authdns-update * 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage * 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage * 11:00 marostegui@dns1004: START - running authdns-update * 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) * 10:51 cmooney@dns3003: END - running authdns-update * 10:49 cmooney@dns3003: START - running authdns-update * 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie * 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie * 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" * 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] * 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json * 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json * 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply * 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply * 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply * 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]]) * 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply * 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply * 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie * 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage * 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie * 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie * 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie * 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage * 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie * 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie * 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash * 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie * 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]]) * 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply * 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage * 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage * 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s) * 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply * 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply * 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment * 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie * 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie * 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash * 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]] * 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] * 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply * 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-02 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-08-01 == * 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-31 == * 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing * 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing * 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing * 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing * 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing * 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad * 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad * 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515 * 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515 * 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage * 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie * 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048 * 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage * 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003" * 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s) * 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment * 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] == 2026-07-30 == * 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts * 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s) * 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts * 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s) * 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]] * 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) * 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment * 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] * 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART * 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie * 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage * 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie * 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie * 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance * 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]] * 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage * 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]] * 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie * 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance * 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json * 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance * 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance * 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie * 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie * 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage * 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance * 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie * 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance * 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance * 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance * 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]] * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage * 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie * 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie * 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" * 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie * 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie * 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json * 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance * 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance * 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance * 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie * 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance * 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) * 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment * 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] * 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie * 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie * 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage * 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) * 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie * 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment * 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance * 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] * 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) * 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance * 14:14 stran@deploy1003: stran: Continuing with deployment * 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json * 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance * 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance * 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie * 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] * 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance * 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) * 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance * 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance * 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance * 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment * 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] * 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance * 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance * 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json * 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance * 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) * 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment * 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]] * 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie * 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance * 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage * 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) * 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] * 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance * 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie * 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json * 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance * 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance * 12:05 ayounsi@dns1004: END - running authdns-update * 12:02 ayounsi@dns1004: START - running authdns-update * 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie * 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance * 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage * 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance * 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance * 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json * 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance * 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie * 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance * 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json * 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance * 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance * 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance * 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie * 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance * 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json * 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance * 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox * 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox * 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance * 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary * 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary * 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json * 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance * 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie * 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance * 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts * 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 * 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json * 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance * 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance * 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance * 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie * 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance * 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie * 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance * 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json * 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance * 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance * 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance * 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance * 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance * 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json * 07:26 klausman@dns2004: END - running authdns-update * 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie * 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie * 07:24 klausman@dns2004: START - running authdns-update * 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet * 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance * 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json * 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance * 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance * 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance * 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json * 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance * 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed * 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json * 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json * 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance * 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json * 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json * 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance * 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance * 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json * 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" * 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance * 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance * 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json * 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] * 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]] == 2026-07-29 == * 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue * 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]] * 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie * 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye * 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]' * 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards * 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s) * 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance * 21:08 aaron@deploy1003: aaron: Continuing with deployment * 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] * 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s) * 20:48 aaron@deploy1003: aaron: Continuing with deployment * 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] * 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance * 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s) * 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json * 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance * 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance * 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment * 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie * 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] * 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94 * 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s) * 19:32 zabe@deploy1003: zabe: Continuing with deployment * 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance * 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] * 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json * 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance * 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance * 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048 * 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048 * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003" * 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance * 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json * 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]]) * 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm * 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json * 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json * 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance * 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json * 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance * 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance * 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance * 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage * 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json * 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet * 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet * 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance * 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance * 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen * 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]] * 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]]) * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022 * 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022 * 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022 * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003" * 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022 * 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm * 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]]) * 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards * 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]] * 17:47 swfrench@dns1004: END - running authdns-update * 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json * 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance * 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance * 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:45 swfrench@dns1004: START - running authdns-update * 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance * 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs * 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance * 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance * 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance * 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json * 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance * 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance * 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s) * 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] * 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s) * 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s) * 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] * 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s) * 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true * 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]] * 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]] * 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s) * 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance * 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json * 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance * 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance * 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance * 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]] * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply * 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance * 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json * 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance * 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance * 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet * 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance * 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment * 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet * 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json * 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance * 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet * 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json * 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance * 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie * 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance * 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance * 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s) * 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] * 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts * 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] * 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm * 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train * 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s) * 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm * 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage * 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]] * 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance * 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] * 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json * 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance * 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance * 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie * 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance * 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie * 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage * 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance * 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json * 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance * 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance * 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance * 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json * 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance * 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance * 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance * 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s) * 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json * 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance * 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance * 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards * 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s) * 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json * 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance * 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm * 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage * 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm * 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie * 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet * 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage * 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage * 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021 * 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021 * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003" * 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance * 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]] * 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage * 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]] * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json * 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance * 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance * 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance * 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye * 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye * 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance * 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage * 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie * 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]]) * 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098'] * 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097'] * 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage * 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json * 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance * 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance * 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors * 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors * 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json * 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance * 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance * 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021 * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015 * 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015 * 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json * 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance * 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s) * 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm * 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) * 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json * 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance * 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage * 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie * 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm * 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285 * 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie * 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002" * 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json * 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) * 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance * 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox * 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance * 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance * 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance * 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage * 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json * 14:03 sukhe@dns1004: END - running authdns-update * 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance * 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance * 14:01 sukhe@dns1004: START - running authdns-update * 14:00 sukhe@dns1004: START - running authdns-update * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance * 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader * 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance * 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json * 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance * 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance * 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply * 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply * 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie * 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm * 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) * 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage * 13:35 stran@deploy1003: stran: Continuing with deployment * 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet * 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] * 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad * 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage * 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) * 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet * 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage * 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265 * 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie * 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet * 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment * 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad * 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad * 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] * 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json * 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance * 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet * 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance * 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s) * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance * 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance * 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment * 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance * 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] * 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json * 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance * 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json * 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie * 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance * 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance * 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003" * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie * 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance * 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json * 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance * 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance * 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json * 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance * 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance * 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]] * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]] * 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]] * 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 12:38 ayounsi@dns1004: END - running authdns-update * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 12:35 ayounsi@dns1004: START - running authdns-update * 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance * 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance * 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json * 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance * 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance * 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts * 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json * 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance * 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance * 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance * 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync * 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance * 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]] * 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance * 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json * 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance * 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance * 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance * 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json * 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance * 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts * 12:00 marostegui: Rename tables [[phab:T425074|T425074]] * 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance * 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync * 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet * 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync * 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync * 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance * 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance * 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json * 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance * 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance * 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json * 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance * 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance * 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json * 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance * 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance * 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]] * 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance * 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]]) * 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json * 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance * 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance * 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json * 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance * 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json * 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance * 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance * 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance * 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json * 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance * 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]] * 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]] * 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance * 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance * 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance * 09:44 btullis@dns1004: END - running authdns-update * 09:42 btullis@dns1004: START - running authdns-update * 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance * 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance * 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet * 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance * 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json * 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance * 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json * 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance * 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance * 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8 * 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5 * 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]] * 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]] * 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org * 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply * 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade * 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org * 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org * 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org * 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org * 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org * 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]] * 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance * 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance * 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade * 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json * 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance * 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]] * 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json * 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance * 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance * 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json * 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance * 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance * 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json * 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance * 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance * 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes * 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts * 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage * 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie * 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance * 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json * 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json * 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance * 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance * 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json * 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance * 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet * 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet * 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage * 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie * 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage * 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie * 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage * 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie * 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage == 2026-07-28 == * 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet * 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet * 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie * 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s) * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards * 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards * 20:54 arlolra@deploy1003: arlolra: Continuing with deployment * 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s) * 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s) * 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] * 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s) * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply * 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:30 arlolra@deploy1003: arlolra: Continuing with deployment * 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm * 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]] * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s) * 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm * 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] * 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage * 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie * 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014 * 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014 * 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm * 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 * 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013 * 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013 * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003" * 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox * 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance * 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json * 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance * 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie * 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" * 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts * 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s) * 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging] * 18:29 sukhe@dns1004: END - running authdns-update * 18:27 sukhe@dns1004: START - running authdns-update * 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance * 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098 * 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098 * 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097 * 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097 * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002" * 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie * 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage * 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie * 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:41 sukhe@dns1004: END - running authdns-update * 17:39 sukhe@dns1004: START - running authdns-update * 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid * 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid * 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003" * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance * 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance * 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance * 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005 * 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005 * 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox * 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad * 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet * 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet * 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003 * 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage * 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet * 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet * 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet * 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet * 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie] * 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts * 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138 * 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138 * 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138 * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003" * 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet * 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json * 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet * 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet * 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet * 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet * 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet * 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet * 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet * 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet * 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm * 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet * 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet * 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet * 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet * 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet * 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox * 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance * 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet * 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet * 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet * 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138 * 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie * 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet * 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance * 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance * 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet * 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet * 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet * 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply * 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet * 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts * 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013 * 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet * 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet * 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet * 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm * 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage * 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet * 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet * 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet * 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet * 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet * 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet * 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s) * 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] * 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s) * 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] * 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards * 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet * 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet * 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet * 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment * 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet * 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment * 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet * 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet * 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance * 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment * 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm * 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet * 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance * 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance * 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance * 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance * 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet * 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet * 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet * 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts * 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet * 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json * 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance * 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance * 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet * 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet * 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance * 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance * 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json * 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance * 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance * 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet * 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet * 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet * 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet * 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet * 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'. * 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'. * 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'. * 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance * 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]] * 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'. * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad * 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json * 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance * 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance * 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]]) * 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]] * 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]]) * 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json * 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance * 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]] * 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance * 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance * 13:45 sukhe: restart pybal on A:lvs-codfw * 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage * 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance * 13:44 btullis@dns1004: END - running authdns-update * 13:42 sukhe: restart pybal on lvs2014 * 13:42 btullis@dns1004: START - running authdns-update * 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json * 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance * 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance * 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json * 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance * 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance * 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]] * 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]] * 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]] * 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards * 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]] * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012 * 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012 * 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]] * 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s) * 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]] * 13:18 swfrench@dns1004: END - running authdns-update * 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment * 13:16 swfrench@dns1004: START - running authdns-update * 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012 * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003 * 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] * 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003" * 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s) * 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s) * 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance * 13:06 esanders@deploy1003: esanders: Continuing with deployment * 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance * 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] * 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json * 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance * 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance * 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance * 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance * 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json * 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw * 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet * 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003" * 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet * 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json * 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance * 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json * 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance * 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance * 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet * 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet * 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet * 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet * 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet * 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet * 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin * 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet * 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin * 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance * 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet * 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet * 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet * 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance * 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json * 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet * 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance * 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance * 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance * 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance * 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance * 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json * 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance * 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance * 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet * 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet * 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet * 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet * 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137 * 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137 * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet * 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet * 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet * 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet * 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance * 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet * 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance * 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance * 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json * 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance * 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" * 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet * 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet * 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet * 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox * 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet * 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw * 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie * 10:28 jforrester@deploy1003: jforrester: Continuing with deployment * 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet * 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker * 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad * 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad * 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet * 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet * 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw * 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw * 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw * 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw * 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance * 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw * 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance * 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet * 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance * 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]] * 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade * 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]] * 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json * 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance * 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet * 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet * 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet * 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet * 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]] * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet * 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker * 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]] * 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]] * 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance * 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance * 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance * 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance * 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json * 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance * 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance * 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json * 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance * 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade * 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance * 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]] * 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json * 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance * 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance * 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json * 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance * 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json * 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance * 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance * 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json * 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance * 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]] * 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]] * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) * 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie * 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" * 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage * 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie * 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie * 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin * 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" * 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox * 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]]) * 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie == 2026-07-27 == * 23:50 Amir1: mass deleting vp8 transcodes * 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye * 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply * 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply * 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]] * 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm * 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye * 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]] * 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]] * 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage * 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]] * 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012 * 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm * 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022 * 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022 * 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022 * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm * 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED * 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004 * 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004 * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003" * 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox * 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards * 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022 * 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm * 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards * 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]] * 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance * 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s) * 20:11 sbisson@deploy1003: sbisson: Continuing with deployment * 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] * 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance * 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s) * 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json * 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance * 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance * 19:23 krinkle@deploy1003: krinkle: Continuing with deployment * 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] * 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance * 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance * 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply * 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm * 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm * 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance * 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s) * 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json * 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance * 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment * 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance * 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] * 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json * 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance * 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance * 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage * 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage * 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance * 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json * 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance * 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance * 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011 * 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011 * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003" * 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011 * 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021 * 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021 * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm * 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance * 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json * 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance * 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main * 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance * 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s) * 17:27 taavi@deploy1003: taavi: Continuing with deployment * 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json * 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance * 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance * 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] * 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance * 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json * 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance * 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance * 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance * 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json * 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance * 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance * 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance * 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance * 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json * 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance * 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance * 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance * 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json * 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance * 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance * 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance * 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance * 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json * 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance * 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance * 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance * 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json * 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance * 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json * 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance * 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards * 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s) * 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance * 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance * 15:18 zabe@deploy1003: zabe: Continuing with deployment * 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] * 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s) * 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm * 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance * 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json * 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance * 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json * 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance * 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm * 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance * 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage * 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010 * 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010 * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003" * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003 * 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply * 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply * 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage * 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org * 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance * 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance * 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance * 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts * 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance * 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts * 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet * 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet * 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance * 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json * 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance * 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010 * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json * 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance * 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020 * 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020 * 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json * 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance * 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm * 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance * 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm * 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts * 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts * 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance * 13:27 Lucas_WMDE: UTC afternoon backport+config window doen * 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s) * 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment * 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]] * 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] * 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]] * 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance * 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . * 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json * 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance * 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance * 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors * 12:30 sukhe@dns1004: END - running authdns-update * 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s) * 12:28 sukhe@dns1004: START - running authdns-update * 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] * 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]] * 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts * 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts * 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance * 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]] * 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade * 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s) * 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json * 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance * 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment * 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] * 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]] * 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing * 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing * 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing * 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing * 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool * 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]] * 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance * 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json * 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json * 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing * 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing * 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json * 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance * 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance * 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008 * 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie * 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away * 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance * 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json * 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance * 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage * 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136 * 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136 * 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136 * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003" * 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox * 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136 * 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie * 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]] * 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet * 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet * 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet * 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json * 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance * 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance * 07:44 phuedx: UTC morning backport window done * 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s) * 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance * 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]] * 07:25 phuedx@deploy1003: phuedx: Continuing with deployment * 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]] * 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json * 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance * 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete * 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] * 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet * 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet * 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning * 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6 * 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4 * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards * 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-26 == * 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm * 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm * 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage * 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage * 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019 * 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018 * 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018 * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm * 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm == 2026-07-25 == * 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering * 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet * 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet * 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards * 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards * 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm * 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm * 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage * 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017 * 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017 * 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024 * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017 * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003" * 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017 * 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024 * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm * 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm * 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses * 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet * 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet * 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards * 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards * 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s) * 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s) * 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host * 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm * 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023 * 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023 * 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023 * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003" * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm * 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage * 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox * 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016 * 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023 * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm * 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm * 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday == 2026-07-24 == * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet * 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet * 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie * 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage * 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet * 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet * 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie * 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts * 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts * 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad * 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage * 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 * 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 * 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage * 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s) * 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135 * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003" * 14:29 krinkle@deploy1003: krinkle: Continuing with deployment * 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm * 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox * 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135 * 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie * 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet * 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003" * 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors * 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors * 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm * 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:30 papaul: reboot mr1-eqsin for maintenance * 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet * 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet * 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts * 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts * 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts * 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie * 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts * 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie * 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie * 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts * 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet * 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet * 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet * 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet * 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet * 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet * 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet * 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet * 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet * 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet * 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet * 09:30 brouberol@dns1004: END - running authdns-update * 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet * 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet * 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts * 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet * 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet * 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts * 09:26 brouberol@dns1004: START - running authdns-update * 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts * 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet * 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet * 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet * 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts * 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet * 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts * 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet * 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet * 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet * 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet * 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet * 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet * 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning * 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4 * 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6 * 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4 * 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue * 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue * 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s) * 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-23 == * 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh * 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003 * 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards * 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s) * 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003 * 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s) * 20:13 dani@deploy1003: dani: Continuing with deployment * 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002" * 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm * 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99) * 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage * 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards * 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s) * 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014 * 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014 * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003" * 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance * 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance * 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance * 17:29 cmooney@dns3003: END - running authdns-update * 17:27 cmooney@dns3003: START - running authdns-update * 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003" * 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance * 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance * 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance * 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json * 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json * 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]] * 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json * 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]] * 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128 * 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance * 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance * 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet * 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie * 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance * 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage * 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072 * 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072 * 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie * 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance * 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s) * 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance * 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment * 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] * 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet * 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003" * 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:19 cmooney@dns2004: END - running authdns-update * 15:17 cmooney@dns2004: START - running authdns-update * 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org * 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet * 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie * 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm * 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance * 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet * 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance * 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet * 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet * 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled * 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet * 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage * 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]] * 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage * 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage * 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance * 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie * 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071 * 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071 * 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards * 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance * 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071 * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003" * 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance * 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance * 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance * 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance * 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014 * 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance * 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm * 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet * 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet * 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008 * 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008 * 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org * 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm * 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]] * 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade * 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade * 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie * 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox * 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071 * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie * 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet * 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet * 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]] * 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003" * 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json * 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s) * 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts * 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment * 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] * 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) * 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies * 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service * 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service * 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling * 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance * 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json * 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance * 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance * 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling * 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling * 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie * 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance * 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet * 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance * 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance * 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 * 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375 * 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad * 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad * 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad * 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad * 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad * 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad * 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie * 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading * 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance * 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing * 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing * 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing * 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing * 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json * 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance * 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing * 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie * 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts * 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage * 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing * 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing * 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing * 10:02 marostegui@dns1004: END - running authdns-update * 10:00 marostegui@dns1004: START - running authdns-update * 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070 * 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070 * 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing * 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing * 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing * 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070 * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003" * 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts * 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070 * 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie * 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet * 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet * 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing * 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing * 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing * 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing * 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing * 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet * 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet * 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage * 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069 * 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069 * 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s) * 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069 * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003" * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 07:24 jiji@deploy1003: jiji: Continuing with deployment * 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules * 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7 * 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2 * 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069 * 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie * 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet * 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet * 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts * 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet * 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet * 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards == 2026-07-22 == * 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003 * 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot` * 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s) * 21:44 sbassett@deploy1003: sbassett: Continuing with deployment * 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] * 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s) * 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert) * 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment * 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing * 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] * 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm * 19:40 mutante: gerrit - one more service restart is needed - restarting * 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing * 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing * 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s) * 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage * 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards * 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016 * 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016 * 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm * 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s) * 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s) * 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] * 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s) * 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated) * 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] * 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s) * 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] * 17:56 Raine: deployment server switchover => deploy1003 is primary now * 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s) * 17:54 mutante: restarting gerrit for maintenance * 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm * 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s) * 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 17:04 Raine: point deployment.eqiad.wmnet to deploy1003 * 17:04 kamila@dns7001: END - running authdns-update * 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 17:02 kamila@dns7001: START - running authdns-update * 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage * 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s) * 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] * 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s) * 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013 * 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013 * 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013 * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003" * 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013 * 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm * 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards * 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s) * 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply * 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply * 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply * 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm * 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s) * 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] * 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s) * 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply * 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006 * 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006 * 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet * 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage * 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm * 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s) * 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] * 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013 * 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]] * 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014 * 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal' * 14:24 sukhe: restart pybal on lvs1019 * 14:24 sukhe: restart pybal on lvs1020 * 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts * 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox * 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s) * 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet * 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003" * 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s) * 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test * 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment * 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] * 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 13:26 sukhe@dns1004: END - running authdns-update * 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 13:24 sukhe@dns1004: START - running authdns-update * 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s) * 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 13:18 stran@deploy2003: stran: Continuing with deployment * 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test * 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test * 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test * 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts * 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003" * 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003 * 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet * 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie * 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s) * 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage * 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates * 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates * 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply * 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply * 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068 * 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068 * 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068 * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003" * 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068 * 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie * 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage * 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet * 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet * 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet * 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s) * 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment * 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet * 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]] * 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates * 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] * 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply * 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates * 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates * 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply * 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s) * 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply * 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply * 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply * 10:50 zabe@deploy2003: zabe: Continuing with deployment * 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates * 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates * 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s) * 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie * 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] * 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage * 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie * 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet * 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet * 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]] * 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003 * 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003" * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates * 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates * 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates * 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning * 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]] * 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie * 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage * 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s) * 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie * 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] * 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s) * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet * 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet * 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment * 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]] * 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade * 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1 * 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade * 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade * 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1 * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply * 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply * 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply * 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply * 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 07:47 phuedx: End of UTC morning backport window * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s) * 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 07:39 phuedx@deploy2003: phuedx: Continuing with deployment * 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003 * 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003" * 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation * 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5 * 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5 * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-21 == * 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive * 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s) * 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024'] * 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards * 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s) * 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm * 20:52 krinkle@deploy2003: krinkle: Continuing with deployment * 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] * 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s) * 20:44 krinkle@deploy2003: krinkle: Rolling back deployment * 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet * 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s) * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet * 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet * 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:29 dani@deploy2003: dani: Continuing with deployment * 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet * 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet * 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet * 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet * 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet * 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet * 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s) * 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 20:10 sbisson@deploy2003: sbisson: Continuing with deployment * 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] * 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s) * 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]]) * 19:59 zabe@deploy2003: zabe: Continuing with deployment * 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] * 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s) * 19:48 zabe@deploy2003: zabe: Continuing with deployment * 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024'] * 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024'] * 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery * 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation * 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad * 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet * 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024'] * 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet * 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003" * 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox * 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm * 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm * 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s) * 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates * 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet * 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet * 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet * 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet * 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet * 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet * 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet * 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates * 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet * 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet * 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.* * 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet * 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie * 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024 * 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024 * 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm * 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance * 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage * 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage * 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance * 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance * 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie * 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm * 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage * 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm * 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]] * 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.* * 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]] * 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance * 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance * 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update * 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s) * 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply * 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update * 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad * 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad * 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance * 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance * 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet * 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet * 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet * 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet * 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts * 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts * 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s) * 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade * 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] * 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]] * 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet * 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update * 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance * 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance * 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance * 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet * 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade * 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json * 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet * 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad * 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply * 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet * 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet * 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet * 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet * 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki * 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade * 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade * 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage * 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json * 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]] * 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet * 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply * 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply * 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply * 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet * 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply * 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet * 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json * 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012 * 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm * 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json * 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance * 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json * 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s) * 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 13:18 kharlan@deploy2003: kharlan: Continuing with deployment * 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json * 13:17 brouberol@dns1004: END - running authdns-update * 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:15 brouberol@dns1004: START - running authdns-update * 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] * 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json * 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards * 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json * 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json * 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json * 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json * 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance * 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json * 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]] * 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json * 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json * 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts * 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json * 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json * 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json * 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance * 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json * 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json * 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json * 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance * 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json * 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json * 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance * 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json * 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json * 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json * 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json * 11:21 XioNoX: put eqiad-drmrs Arelion link in service * 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json * 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json * 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json * 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance * 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json * 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json * 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance * 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json * 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json * 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json * 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json * 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance * 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance * 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json * 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json * 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance * 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel * 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet * 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet * 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet * 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json * 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json * 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]] * 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]] * 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003 * 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s) * 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment * 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply * 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org * 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] * 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply * 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org * 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades * 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades * 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync * 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync * 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync * 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync * 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync * 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync * 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage * 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003 * 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003" * 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning * 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet * 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet * 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s) * 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s) * 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] * 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm == 2026-07-20 == * 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis * 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s) * 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] * 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 22:39 maryum: Deployed security fixes for several security bugs * 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards * 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]] * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s) * 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw * 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage * 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service * 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts * 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service * 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw * 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s) * 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007 * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007 * 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007 * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003" * 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json * 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s) * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet * 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet * 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet * 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment * 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json * 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] * 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json * 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007 * 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm * 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json * 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance * 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json * 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json * 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards * 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards * 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm * 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json * 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json * 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance * 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json * 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json * 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json * 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json * 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json * 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance * 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json * 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s) * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020'] * 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json * 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json * 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json * 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw * 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet * 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw * 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet * 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance * 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance * 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host) * 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet * 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet * 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020'] * 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet * 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet * 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet * 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json * 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020'] * 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage * 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm * 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet * 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet * 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet * 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json * 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s) * 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet * 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011 * 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm * 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json * 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json * 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s) * 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json * 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance * 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json * 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json * 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet * 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet * 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json * 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json * 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json * 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance * 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020 * 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020 * 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020 * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json * 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002" * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet * 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json * 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox * 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020 * 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm * 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json * 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json * 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance * 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json * 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet * 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json * 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet * 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json * 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023 * 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm * 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json * 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) * 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json * 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance * 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment * 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s) * 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards * 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet * 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet * 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm * 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica * 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] * 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] * 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards * 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage * 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet * 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet * 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]] * 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019 * 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019 * 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet * 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet * 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet * 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie * 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage * 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet * 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet * 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet * 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet * 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019 * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003" * 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet * 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019 * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm * 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027 * 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm * 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw * 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet * 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet * 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage * 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage * 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie * 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie * 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts * 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts * 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts * 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts * 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts * 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie * 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie * 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts * 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts * 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet * 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet * 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet * 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage * 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage * 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet * 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie * 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts * 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts * 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts * 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie * 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage * 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage * 10:06 blake@deploy2003: Stopping before sync operations * 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie * 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie * 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts * 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts * 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts * 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage * 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie * 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie * 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s) * 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values * 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage * 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie * 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie * 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie * 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie * 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie * 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-18 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards == 2026-07-17 == * 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards * 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards * 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm * 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm * 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018 * 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018 * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003" * 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018 * 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm * 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026 * 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards * 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm * 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards * 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s) * 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s) * 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s) * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s) * 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s) * 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage * 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage * 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s) * 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm * 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017 * 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017 * 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm * 18:30 bking@dns1004: END - running authdns-update * 18:28 bking@dns1004: START - running authdns-update * 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet * 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet * 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:46 dzahn@dns1006: END - running authdns-update * 17:44 dzahn@dns1006: START - running authdns-update * 17:44 dzahn@dns1006: END - running authdns-update * 17:42 dzahn@dns1006: START - running authdns-update * 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s) * 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment * 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] * 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage * 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]] * 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet * 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet * 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339 * 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339 * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003" * 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339 * 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie * 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet * 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet * 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet * 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet * 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 13:24 blake@dns1004: END - running authdns-update * 13:22 blake@dns1004: START - running authdns-update * 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet * 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet * 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication * 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts * 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet * 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet * 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts * 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet * 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet * 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet * 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet * 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet * 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie * 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage * 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338 * 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338 * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003" * 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338 * 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie * 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet * 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet * 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet * 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet * 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage * 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet * 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337 * 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337 * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003" * 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337 * 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie * 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet * 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet * 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet * 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet * 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet * 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.) * 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.) * 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.) * 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.) * 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1 * 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn * 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage * 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336 * 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336 * 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336 * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003" * 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336 * 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie * 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet * 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet * 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet * 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet * 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet * 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet * 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet * 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet * 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet * 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335 * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003" * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335 * 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie * 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet * 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet * 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99) * 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet * 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet * 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet * 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet * 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet * 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster * 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges * 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet * 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet * 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards * 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply * 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply * 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply * 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards * 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet * 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm * 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm * 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie == 2026-07-16 == * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage * 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage * 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage * 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 * 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025 * 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 * 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm * 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" * 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm * 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 * 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie * 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet * 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet * 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there * 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad * 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet * 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet * 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards * 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:56 Amir1: deleting echo notifications from 2015 in group0 * 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm * 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s) * 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] * 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet * 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet * 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage * 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) * 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment * 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 * 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 * 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" * 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage * 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 * 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie * 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet * 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet * 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet * 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie * 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) * 21:40 sbassett@deploy2003: sbassett: Continuing with deployment * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 * 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 * 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] * 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" * 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) * 21:26 sbassett@deploy2003: sbassett: Continuing with deployment * 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie * 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] * 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025 * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm * 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage * 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage * 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 * 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 * 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" * 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet * 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 * 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie * 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet * 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet * 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 * 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 * 20:36 aude@deploy2003: aude: Continuing with deployment * 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] * 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down * 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply * 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply * 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply * 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply * 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply * 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply * 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply * 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" * 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. * 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json * 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json * 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]] * 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet * 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json * 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]] * 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet * 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet * 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet * 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply * 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet * 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet * 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply * 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply * 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply * 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply * 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply * 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply * 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet * 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet * 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply * 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply * 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply * 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply * 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet * 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet * 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet * 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet * 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox * 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet * 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 * 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie * 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet * 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet * 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s) * 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet * 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet * 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards * 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) * 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s) * 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage * 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 * 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 * 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie * 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet * 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet * 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) * 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" * 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 * 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie * 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] * 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage * 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet * 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet * 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet * 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet * 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet * 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" * 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) * 17:04 jiji@deploy2003: jiji: Continuing with deployment * 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] * 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie * 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 * 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" * 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin * 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet * 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet * 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet * 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet * 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet * 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad * 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet * 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox * 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 * 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage * 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie * 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet * 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet * 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet * 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet * 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie * 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 * 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie * 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet * 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin * 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet * 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet * 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet * 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) * 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 * 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 * 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet * 16:07 jiji@deploy2003: jiji: Continuing with deployment * 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox * 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet * 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 * 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie * 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet * 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet * 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet * 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet * 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet * 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie * 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage * 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet * 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) * 15:37 jiji@deploy2003: jiji: Continuing with deployment * 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 * 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 * 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet * 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet * 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 * 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet * 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet * 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm * 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" * 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 * 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie * 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad * 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet * 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet * 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 * 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]] * 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 * 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" * 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 * 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet * 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 * 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie * 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie * 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet * 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet * 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet * 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage * 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet * 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet * 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet * 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet * 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet * 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet * 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet * 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet * 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet * 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet * 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet * 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet * 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet * 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s) * 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. * 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet * 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet * 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet * 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet * 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s) * 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]] * 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] * 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm * 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade * 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet * 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet * 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade * 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 * 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 * 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 * 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 * 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 * 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet * 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet * 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet * 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet * 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]] * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet * 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet * 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet * 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw * 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet * 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet * 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet * 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet * 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage * 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad * 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet * 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org * 13:27 sukhe@dns1004: END - running authdns-update * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet * 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet * 13:25 sukhe@dns1004: START - running authdns-update * 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet * 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 * 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 * 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 * 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet * 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" * 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet * 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet * 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org * 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.* * 13:14 cdobbins@dns1004: END - running authdns-update * 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) * 13:13 cdobbins@dns1004: START - running authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox * 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 * 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie * 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet * 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet * 13:09 sbisson@deploy2003: sbisson: Continuing with deployment * 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] * 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org * 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet * 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet * 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet * 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet * 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet * 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet * 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet * 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet * 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet * 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet * 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org * 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet * 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet * 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org * 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet * 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw * 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet * 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet * 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet * 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org * 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox) * 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad * 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams * 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin * 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin * 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad * 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw * 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet * 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad * 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw * 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie * 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet * 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet * 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie * 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) * 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet * 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage * 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie * 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet * 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage * 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage * 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet * 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet * 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet * 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie * 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie * 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie * 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet * 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover * 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie * 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:57 tappof: bump space for prometheus k8s-dse in eqiad * 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet * 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet * 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet * 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet * 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet * 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage * 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet * 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage * 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie * 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie * 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet * 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad * 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie * 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover * 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover * 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json * 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json * 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]] * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json * 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]] * 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad * 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet * 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage * 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet * 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet * 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet * 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet * 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet * 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet * 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org * 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie * 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org * 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet * 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet * 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet * 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet * 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet * 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet * 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet * 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet * 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet * 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet * 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet * 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet * 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet * 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet * 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet * 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet * 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet * 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet * 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet * 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet * 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet * 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet * 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet * 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet * 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet * 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet * 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet * 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet * 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards * 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet * 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet * 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet * 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet * 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s) * 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage * 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) * 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet * 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet * 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet * 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet * 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet * 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet * 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet * 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet * 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet == 2026-07-15 == * 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet * 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet * 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet * 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet * 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet * 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet * 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm * 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet * 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage * 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet * 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet * 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet * 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet * 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie * 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet * 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet * 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet * 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet * 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet * 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm * 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages * 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet * 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet * 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet * 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet * 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002'] * 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm * 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet * 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet * 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002'] * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002'] * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet * 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet * 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons. * 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet * 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad * 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun! * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001'] * 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001'] * 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm * 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch * 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet * 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet * 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie * 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage * 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262 * 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262 * 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262 * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003" * 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox * 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262 * 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet * 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie * 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet * 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet * 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet * 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]] * 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad * 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad * 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo * 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet * 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs * 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet * 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo * 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet * 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs * 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update * 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]] * 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet * 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet * 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply * 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply * 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet * 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw * 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet * 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update * 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet * 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet * 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet * 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply * 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:06 sukhe: sre.dns.roll-reboot to resume later * 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org * 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update * 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply * 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet * 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply * 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet * 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet * 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet * 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet * 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply * 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet * 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet * 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw * 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply * 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply * 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply * 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet * 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet * 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org * 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet * 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet * 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet * 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet * 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet * 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie * 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet * 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet * 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org * 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet * 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet * 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply * 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet * 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet * 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie * 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet * 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet * 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet * 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet * 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet * 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet * 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet * 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org * 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet * 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet * 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet * 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage * 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet * 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet * 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw * 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet * 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org * 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet * 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie * 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet * 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons. * 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet * 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet * 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes * 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie * 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org * 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad * 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie * 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet * 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet * 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet * 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org * 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad * 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw * 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet * 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet * 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet * 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet * 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet * 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet * 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet * 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw * 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org * 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet * 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet * 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet * 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet * 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage * 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]] * 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet * 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet * 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet * 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad * 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie * 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie * 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org * 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad * 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet * 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet * 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org * 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm * 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet * 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet * 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet * 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install * 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad * 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes * 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]] * 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org * 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad * 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet * 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org * 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet * 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet * 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet * 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm * 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001 * 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001 * 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet * 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org * 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install * 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet * 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001 * 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003" * 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs * 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs * 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install * 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad * 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad * 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org * 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox) * 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo * 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet * 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001 * 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie * 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards * 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s) * 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet * 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet * 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet * 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet * 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage * 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw * 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet * 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009 * 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009 * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003" * 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009 * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie * 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet * 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet * 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet * 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet * 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet * 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw * 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes * 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s) * 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie * 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet * 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet * 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates * 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates * 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] * 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet * 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates * 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache * 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates * 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie * 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet * 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad * 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010 * 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010 * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003" * 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010 * 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie * 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet * 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet * 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet * 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet * 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage * 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet * 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates * 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates * 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet * 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet * 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s) * 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet * 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet * 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie * 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie * 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet * 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet * 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet * 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet * 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet * 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet * 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet * 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates * 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache * 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates * 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage * 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet * 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet * 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet * 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet * 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet * 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie * 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] * 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet * 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]] * 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet * 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011 * 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011 * 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet * 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011 * 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003" * 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet * 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011 * 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie * 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet * 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet * 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet * 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage * 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet * 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet * 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet * 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad * 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad * 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie * 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie * 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage * 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage * 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates * 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012 * 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates * 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012 * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003" * 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie * 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012 * 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie * 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet * 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet * 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet * 08:30 elukey@dns1004: END - running authdns-update * 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:27 elukey@dns1004: START - running authdns-update * 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet * 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet * 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org * 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie * 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org * 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates * 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie * 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org * 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org * 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org * 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie * 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org * 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage * 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org * 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013 * 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013 * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013 * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003" * 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie * 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013 * 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie * 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s) * 07:09 kharlan@deploy2003: kharlan: Continuing with deployment * 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image * 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply == 2026-07-14 == * 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru * 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet * 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru * 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet * 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet * 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet * 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet * 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet * 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot * 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance * 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot * 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet * 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet * 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie * 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) * 20:24 sbassett@deploy2003: sbassett: Continuing with deployment * 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] * 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage * 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) * 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment * 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet * 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] * 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie * 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet * 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet * 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet * 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet * 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet * 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) * 19:07 jforrester@deploy2003: jforrester: Continuing with deployment * 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] * 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet * 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator) * 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) * 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet * 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet * 17:32 swfrench@deploy2003: swfrench: Continuing with deployment * 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance * 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance * 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image * 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet * 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia * 16:50 sukhe: pool cp2046 * 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet * 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent" * 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot * 16:32 mutante: contint1003 - main CI server - rebooting * 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance * 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance * 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet * 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet * 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe * 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet * 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie * 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie * 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 * 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 * 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" * 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014 * 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie * 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie * 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts * 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance * 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance * 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s) * 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie * 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]]) * 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s) * 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie * 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment * 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] * 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply * 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply * 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply * 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply * 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage * 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage * 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]] * 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie * 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie * 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance * 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie * 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie * 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw * 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync * 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet * 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet * 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync * 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org * 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync * 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync * 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org * 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet * 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet * 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw * 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw * 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir * 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync * 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync * 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage * 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough * 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy * 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum * 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum * 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet * 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org * 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie * 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]] * 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie * 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet * 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet * 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet * 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie * 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet * 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org * 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet * 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet * 13:18 elukey@dns1004: END - running authdns-update * 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet * 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet * 13:16 elukey@dns1004: START - running authdns-update * 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 * 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet * 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet * 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet * 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet * 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet * 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet * 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet * 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org * 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage * 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet * 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica * 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance * 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance * 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance * 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org * 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance * 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance * 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance * 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet * 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw * 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance * 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage * 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw * 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw * 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade * 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw * 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet * 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw * 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie * 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet * 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie * 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]] * 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org * 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) * 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet * 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet * 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum * 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet * 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy * 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir * 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough * 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie * 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org * 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie * 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet * 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet * 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet * 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org * 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org * 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet * 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage * 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie * 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org * 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org * 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet * 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet * 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs * 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org * 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org * 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]]) * 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet * 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org * 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet * 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet * 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet * 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage * 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage * 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet * 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet * 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 * 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage * 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet * 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet * 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet * 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet * 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet * 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet * 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet * 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage * 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" * 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet * 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet * 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 * 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 * 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet * 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet * 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet * 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync * 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie * 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync * 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 * 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie * 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie * 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet * 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad) * 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync * 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync * 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment * 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage * 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage * 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] * 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet * 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie * 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie * 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie * 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" * 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) * 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie * 11:03 kharlan@deploy2003: kharlan: Continuing with deployment * 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox * 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] * 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) * 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:52 marostegui@dns1004: START - running authdns-update * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage * 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage * 10:43 kharlan@deploy2003: kharlan: Continuing with deployment * 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie * 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover * 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 * 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie * 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet * 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet * 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet * 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet * 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie * 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] * 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie * 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert * 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage * 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage * 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 * 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage * 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie * 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie * 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover * 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie * 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie * 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 * 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 * 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage * 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage * 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet * 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet * 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet * 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) * 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] * 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie * 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie * 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) * 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie * 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie * 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] * 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) * 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" * 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] * 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json * 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox * 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 * 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie * 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet * 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet * 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json * 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]] * 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json * 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]] * 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage * 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:29 marostegui@dns1004: END - running authdns-update * 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet * 08:27 marostegui@dns1004: START - running authdns-update * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet * 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet * 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie * 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet * 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet * 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet * 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet * 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie * 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie * 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet * 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet * 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet * 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet * 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet * 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet * 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet * 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet * 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors * 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet * 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors * 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki) * 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage * 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet * 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet * 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet * 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors * 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors * 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw) * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet * 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet * 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie * 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie * 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet * 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw) * 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org * 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org * 06:25 marostegui@dns1004: END - running authdns-update * 06:23 marostegui@dns1004: START - running authdns-update * 06:22 marostegui@dns1004: END - running authdns-update * 06:20 marostegui@dns1004: START - running authdns-update * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot * 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 05:26 marostegui@dns1004: END - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:24 marostegui@dns1004: START - running authdns-update * 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org * 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org * 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) * 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s) * 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-13 == * 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet * 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet * 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs * 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet * 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet * 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs * 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]] * 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]] * 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]] * 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s) * 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]] * 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts * 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s) * 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) * 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment * 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there * 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] * 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts * 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie * 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet * 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet * 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs * 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet * 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet * 17:06 dzahn@dns1006: END - running authdns-update * 17:04 dzahn@dns1006: START - running authdns-update * 17:01 dzahn@dns1006: END - running authdns-update * 16:59 dzahn@dns1006: START - running authdns-update * 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s) * 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]] * 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]] * 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]]) * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts * 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs * 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s) * 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync * 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync * 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync * 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync * 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync * 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync * 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync * 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 * 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 * 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . * 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet * 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet * 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet * 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet * 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet * 15:36 sukhe: restart pybal on lvs1020 * 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:28 marostegui@dns1004: END - running authdns-update * 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:27 marostegui@dns1004: START - running authdns-update * 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot * 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie * 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]] * 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 14:02 marostegui@dns1004: END - running authdns-update * 14:00 marostegui@dns1004: START - running authdns-update * 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet * 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet * 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.* * 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage * 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie * 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs * 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) * 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment * 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers * 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] * 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie * 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) * 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment * 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] * 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:23 Msz2001: Deployed changes to private code for Suggested Investigations * 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage * 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) * 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment * 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] * 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) * 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie * 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s) * 11:55 zabe@deploy2003: zabe: Continuing with deployment * 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie * 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] * 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage * 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 * 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 * 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 * 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 * 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply * 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json * 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie * 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie * 09:42 marostegui@dns1004: END - running authdns-update * 09:40 marostegui@dns1004: START - running authdns-update * 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage * 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie * 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply * 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply * 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply * 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet * 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply * 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet * 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 * 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie * 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet * 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet * 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing * 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 * 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet * 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet * 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet * 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet * 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:07 marostegui@dns1004: END - running authdns-update * 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet * 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet * 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 marostegui@dns1004: START - running authdns-update * 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 08:04 marostegui@dns1004: START - running authdns-update * 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org * 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet * 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet * 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet * 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:52 Msz2001: UTC morning backport+config window done * 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie * {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}} * 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet * 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 * 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 * 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment * {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}} * 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing * {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}} * 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) * 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org * 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment * 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org * 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet * 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet * 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org * 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] * 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org * 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org * 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org * 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org * 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]] * 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot * 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot * 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot * 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot * 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot * 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning * 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]] * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-12 == * 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json * 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json * 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]] * 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json * 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]] * 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-11 == * 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-10 == * 19:12 jhathaway@dns1004: END - running authdns-update * 19:10 jhathaway@dns1004: START - running authdns-update * 18:23 mutante: vrts2002 rebooting (not the active host) * 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts) * 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw * 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw * 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]]) * 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet * 17:08 jhathaway@dns1004: END - running authdns-update * 17:07 jhathaway@dns1004: START - running authdns-update * 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown * 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet * 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet * 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet * 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet * 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet * 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet * 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet * 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet * 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet * 16:25 mutante: gitlab-runners (production) rebooting cluster one by one * 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting * 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting * 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance * 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet * 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet * 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet * 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet * 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet * 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet * 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet * 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet * 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet * 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet * 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet * 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet * 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet * 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet * 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet * 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet * 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie * 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet * 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet * 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet * 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet * 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet * 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet * 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet * 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet * 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet * 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet * 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage * 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet * 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet * 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet * 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet * 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org * 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054 * 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054 * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003" * 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org * 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054 * 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie * 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet * 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet * 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie * 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply * 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply * 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet * 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet * 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet * 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet * 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet * 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade * 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie * 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts * 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts * 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie * 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie * 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet * 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw) * 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet * 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw * 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm * 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply * 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm * 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply * 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet * 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw * 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] * 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053 * 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053 * 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053 * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003" * 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox * 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053 * 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie * 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet * 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet * 09:04 brouberol@dns1004: END - running authdns-update * 09:03 brouberol@dns1004: START - running authdns-update * 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s) * 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] * 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s) * 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] * 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s) * 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] * 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}} * 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet * 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet * 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply * 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade * 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply * 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply * 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts * 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts * 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-09 == * 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) * 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment * 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] * 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet * 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet * 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie * 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage * 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 * 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" * 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 * 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie * 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet * 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet * 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:42 maryum: Deploy fix for [[phab:T431684|T431684]] * 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage * 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) * 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002 * 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie * 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s) * 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards * 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment * 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm * 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots * 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] * 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 * 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet * 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) * 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm * 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet * 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet * 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]] * 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage * 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) * 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views * 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie * 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org * 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org * 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply * 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply * 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply * 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply * 17:09 mutante: zuul[12]00[123] - rebooting for maintenance * 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster * 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) * 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster * 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster * 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance * 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster * 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply * 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply * 16:49 mutante: planet1003/planet2003 - rebooting * 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating * 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:51 jynus: restarting backupmon1001 * 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply * 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply * 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply * 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply * 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts * 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts * 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet * 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet * 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet * 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync * 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync * 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie * 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade * 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]] * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply * 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox * 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet * 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet * 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:57 moritzm: installing requests security updates * 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet * 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet * 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing python-cryptography security updates * 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet * 13:44 Msz2001: UTC afternoon config+backport window is done * 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297 * 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org * 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) * 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org * 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 * 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 * 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" * 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] * 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox * 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 * 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie * 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet * 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet * 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) * 13:13 jforrester@deploy2003: jforrester: Continuing with deployment * 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] * 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply * 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply * 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply * 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply * 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply * 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org * 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org * 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org * 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 * 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004 * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org * 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org * 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet * 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet * 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet * 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet * 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet * 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet * 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org * 11:55 jmm@dns1004: END - running authdns-update * 11:53 jmm@dns1004: START - running authdns-update * 11:50 jmm@dns1004: END - running authdns-update * 11:48 jmm@dns1004: START - running authdns-update * 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org * 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet * 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet * 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org * 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org * 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org * 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org * 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org * 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox * 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org * 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet * 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 * 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003 * 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org * 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org * 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org * 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org * 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org * 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]] * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org * 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org * 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet * 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]] * 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]] * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org * 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B * 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org * 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org * 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org * 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage * 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org * 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 * 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org * 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" * 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet * 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet * 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox * 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004 * 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm * 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet * 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet * 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet * 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet * 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox * 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet * 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet * 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet * 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet * 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet * 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet * 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet * 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet * 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet * 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance * 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet * 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance * 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment * 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot * 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 * 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 * 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" * 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003 * 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet * 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance * 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance * 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage * 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:31 hashar@deploy2003: Rolling back deployment * 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]] * 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm * 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet * 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet * 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) * 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet * 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]] * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet * 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet * 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet * 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance * 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet * 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance * 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance * 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 * 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie * 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) * 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment * 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] * 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage * 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie * 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie * 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie * 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]] * 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.* * 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad * 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad * 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) * 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image == 2026-07-08 == * 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]]) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" * 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) * 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment * 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] * 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage * 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie * 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]]) * 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw * 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw * 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw * 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 * 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" * 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002 * 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie * 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage * 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye * 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" * 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]]) * 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) * 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye * 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage * 20:30 cjming@deploy2003: cjming: Continuing with deployment * 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie * 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie * 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] * 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage * 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw * 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw * 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw * 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003) * 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet * 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet * 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) * 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) * 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc * 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw * 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw * 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw * 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster * 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie * 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw * 19:19 mutante: gerrit - replacing private key for registerEmail verification * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw * 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw * 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw * 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw * 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw * 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw * 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw * 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru * 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru * 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru * 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru * 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage * 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie * 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]] * 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'. * 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie * 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage * 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 * 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 * 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. * 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage * 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. * 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox * 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 * 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie * 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'. * 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. * 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'. * 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie * 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s) * 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] * 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s) * 17:42 kamila@deploy2003: Rolling back deployment * 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie * 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) * 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:11 kamila@dns1005: END - running authdns-update * 17:09 kamila@dns1005: START - running authdns-update * 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage * 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet * 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet * 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover * 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) * 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie * 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server * 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie * 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie * 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie * 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage * 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage * 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test * 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test * 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 * 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" * 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie * 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox * 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 * 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie * 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie * 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet * 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet * 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie * 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply * 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply * 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage * 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test * 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test * 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'. * 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'. * 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet * 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet * 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]] * 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]] * 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service" * 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie * 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004 * 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004 * 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]] * 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie * 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test * 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:29 swfrench@dns1004: END - running authdns-update * 14:29 moritzm: installing jackson-core security updates * 14:27 swfrench@dns1004: START - running authdns-update * 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie * 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'. * 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'. * 14:19 moritzm: installing librabbitmq security updates * 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet * 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet * 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test * 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet * 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet * 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie * 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad * 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:54 moritzm: installing libcap2 security updates * 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage * 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) * 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool * 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad * 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet * 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet * 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 * 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 * 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" * 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:39 moritzm: installing krb5 security updates * 13:37 Lucas_WMDE: UTC afternoon backport+config window done * 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie * 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage * 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 * 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie * 13:30 sbisson@deploy1003: sbisson: Continuing with deployment * 13:30 moritzm: installing openssh security updates * 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] * 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 * 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 * 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) * 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage * 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" * 13:17 stran@deploy1003: stran: Continuing with deployment * 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 * 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage * 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie * 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw * 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie * 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 13:04 moritzm: installing jq security updates * 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage * 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw * 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 * 12:43 moritzm: installing Python 3.11 security updates * 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" * 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 * 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie * 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet * 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet * 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL * 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie * 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie * 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie * 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie * 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie * 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet * 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet * 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]] * 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage * 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage * 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie * 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie * 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage * 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie * 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet * 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie * 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test * 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance * 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) * 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit * 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) * 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart * 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance * 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet * 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet * 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet * 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 * 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 * 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 * 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage * 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 * 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]] * 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 * 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 * 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet * 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance * 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet * 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet * 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance * 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance * 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie * 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie * 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance * 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 * 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage * 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 07:29 moritzm: installing gnutls28 security updates * 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie * 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie * 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage * 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm * 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage * 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie * 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie * 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie * 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie * 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage * 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage * 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage * 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie * 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie * 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s) * 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] * 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s) * 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] * 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie * 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage * 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage * 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie * 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie == 2026-07-07 == * 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage * 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) * 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie * 22:07 hashar: Restarting Gerrit on gerrit2003 * 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply * 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply * 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet * 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage * 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet * 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" * 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) * 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" * 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie * 20:17 arlolra@deploy1003: arlolra: Continuing with deployment * 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] * 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space * 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie * 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie * 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie * 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie * 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie * 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.* * 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:17 jasmine@dns1004: END - running authdns-update * 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage * 19:15 jasmine@dns1004: START - running authdns-update * 19:15 cdobbins@dns1004: END - running authdns-update * 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 19:13 cdobbins@dns1004: START - running authdns-update * 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update * 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage * 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie * 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie * 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie * 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]]) * 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]]) * 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]] * 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]] * 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]] * 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]] * 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]] * 17:43 swfrench@dns1004: END - running authdns-update * 17:40 swfrench@dns1004: START - running authdns-update * 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]] * 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet * 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet * 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet * 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet * 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]] * 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet * 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie * 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]] * 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet * 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm * 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]] * 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]] * 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage * 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 * 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794 * 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie * 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie * 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage * 15:54 mutante: jenkins down in planned maintenance window * 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet * 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet * 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie * 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm * 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie * 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage * 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org * 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 * 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 * 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org * 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 * 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" * 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie * 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 * 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie * 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie * 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org * 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org * 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s) * 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] * 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s) * 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] * 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) * 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] * 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) * 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] * 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) * 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] * 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance * 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage * 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table * 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage * 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie * 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad * 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 * 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 * 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) * 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs * 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 * 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] * 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) * 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) * 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication * 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]] * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" * 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw * 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 * 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie * 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet * 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs * 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet * 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet * 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie * 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage * 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet * 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet * 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet * 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw * 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie * 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases * 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 * 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) * 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie * 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment * 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage * 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply * 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage * 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply * 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage * 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply * 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 13:38 moritzm: installing e2fsprogs updates from Trixie point release * 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] * 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie * 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade * 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie * 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie * 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.* * 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie * 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie * 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage * 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]] * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie * 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G * 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:13 jmm@dns1004: END - running authdns-update * 13:12 jmm@dns1004: START - running authdns-update * 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage * 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie * 13:07 jmm@dns1004: END - running authdns-update * 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm * 13:05 jmm@dns1004: START - running authdns-update * 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]] * 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet * 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie * 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet * 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet * 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage * 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage * 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet * 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki> * 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" * 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie * 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage * 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 * 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 * 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie * 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie * 12:30 jmm@dns1004: END - running authdns-update * 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" * 12:28 jmm@dns1004: START - running authdns-update * 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox * 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 * 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet * 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie * 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet * 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet * 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]] * 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter * 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) * 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad * 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet * 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) * 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet * 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet * 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad * 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie * 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence * 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]] * 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin * 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin * 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet * 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie * 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet * 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie * 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie * 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage * 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage * 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie * 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie * 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s) * 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage * 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s) * 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie * 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] * 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 * 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 * 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie * 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 * 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" * 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache * 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot * 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet * 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet * 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox * 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 * 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie * 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage * 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage * 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie * 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates * 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates * 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache * 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates * 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply * 08:45 filippo@dns1006: END - running authdns-update * 08:43 filippo@dns1006: START - running authdns-update * 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]] * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates * 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates * 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet * 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates * 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache * 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates * 07:42 Msz2001: Deployed private patch for Suggested Ivestigations * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates * 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates * 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates * 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache * 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 * 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" * 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet * 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates * 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache * 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:42 moritzm: install nginx security updates * 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates * 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates * 06:19 moritzm: installing php8.2 security updates * 06:15 moritzm: installing php8.4 security updates * 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) * 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s) * 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-06 == * 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s) * 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie * 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment * 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug) * 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] * 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage * 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie * 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes * 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes * 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]] * 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie * 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie * 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage * 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s) * 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment * 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage * 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] * 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie * 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage * 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie * 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage * 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie * 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie * 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie * 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage * 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie * 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage * 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie * 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie * 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie * 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage * 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage * 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage * 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie * 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie * 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie * 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie * 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie * 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie * 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage * 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage * 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage * 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie * 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie * 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts * 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie * 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s) * 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. * 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. * 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. * 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie * 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. * 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. * 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. * 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage * 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage * 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie * 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie * 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage * 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm * 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]] * 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]] * 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet * 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet * 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet * 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97) * 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw * 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet * 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet * 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet * 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet * 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw * 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s) * 12:24 krinkle@deploy1003: krinkle: Continuing with deployment * 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet * 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet * 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] * 11:57 moritzm: installing curl security updates * 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 11:31 moritzm: installing nano security updates * 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]] * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie * 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:50 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]] * 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 elukey: spicerack 13.0.0 deployed on cumin2002 * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie * 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia * 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]] * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply * 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply * 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet * 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]] * 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet * 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s) * 07:57 fabfur: repooled cp4038 * 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:53 moritzm: installing pyjwt security updates * 07:47 moritzm: installing openjpeg2 security updates * 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment * 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:38 moritzm: installing python-urllib3 security updates * 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw * 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure * 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.* * 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.* * 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw * 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] * 06:13 moritzm: installing Linux 6.12.95 on trixie hosts * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 * 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 * 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-05 == * 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-04 == * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-03 == * 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo * 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo * 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]] * 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]] * 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo * 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo * 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 14:40 cmooney@dns3003: END - running authdns-update * 14:26 cmooney@dns3003: START - running authdns-update * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003" * 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 13:38 sukhe@dns1004: END - running authdns-update * 13:35 sukhe@dns1004: START - running authdns-update * 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet * 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]] * 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet * 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet * 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet * 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply * 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply * 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet * 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet * 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003" * 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox * 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet * 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" * 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox * 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet * 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]] * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" * 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox * 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org * 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image == 2026-07-02 == * 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie * 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage * 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie * 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage * 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s) * 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]] * 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003 * 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s) * 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment * 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s) * 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] * 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie * 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha * 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] * 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s) * 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:13 sbassett@deploy1003: sbassett: Continuing with deployment * 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] * 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage * 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards * 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie * 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]] * 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm * 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet * 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet * 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet * 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005" * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply * 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply * 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s) * 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment * 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] * 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage * 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply * 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply * 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply * 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts * 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply * 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply * 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply * 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply * 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003" * 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003 * 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003 * 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003" * 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards * 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:57 tappof: bump space for prometheus k8s-aux in eqiad * 16:55 cmooney@dns3003: END - running authdns-update * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003" * 16:53 cmooney@dns3003: START - running authdns-update * 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure) * 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" * 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 16:39 rzl@dns1004: END - running authdns-update * 16:37 rzl@dns1004: START - running authdns-update * 16:36 rzl@dns1004: START - running authdns-update * 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s) * 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:30 rzl@deploy1003: rzl: Continuing with deployment * 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage * 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]] * 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync * 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync * 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm * 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm * 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet * 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates * 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache * 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates * 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates * 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates * 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm * 15:24 moritzm: installing busybox updates from bookworm point release * 15:20 moritzm: installing busybox updates from trixie point release * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates * 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache * 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates * 15:13 moritzm: installing giflib security updates * 15:08 moritzm: installing Tomcat security updates * 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. * 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. * 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates * 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) * 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache * 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates * 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003 * 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003" * 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json * 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json * 14:32 moritzm: installing libdbi-perl security updates * 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json * 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json * 14:12 moritzm: installing rsync security updates * 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json * 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance * 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 14:06 Tran: Deployed patch for [[phab:T427287|T427287]] * 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover * 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover * 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json * 13:54 moritzm: installing sed security updates * 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json * 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]] * 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json * 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]] * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply * 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply * 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw * 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART * 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough * 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw * 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw * 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw * 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org * 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough * 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough * 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s) * 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment * 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 12:10 jmm@dns1004: END - running authdns-update * 12:07 jmm@dns1004: START - running authdns-update * 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003" * 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox * 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet * 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet * 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet * 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet * 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet * 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) * 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges * 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling * 10:49 jmm@dns1004: END - running authdns-update * 10:47 jmm@dns1004: START - running authdns-update * 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json * 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json * 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet * 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) * 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]] * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" * 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet * 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet * 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling * 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json * 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox * 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet * 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission * 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json * 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json * 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance * 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover * 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover * 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json * 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json * 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]] * 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json * 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json * 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]] * 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json * 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s) * 09:13 moritzm: installing libgcrypt20 security updates * 09:12 kharlan@deploy1003: kharlan: Continuing with deployment * 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json * 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] * 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s) * 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json * 08:57 kharlan@deploy1003: kharlan: Continuing with deployment * 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json * 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance * 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s) * 08:21 cscott@deploy1003: cscott: Continuing with deployment * 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed * 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . * 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s) * 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . * 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org * 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:57 cscott@deploy1003: cscott: Continuing with deployment * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . * 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org * 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply * 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org * 07:44 moritzm: installing node-lodash security updates * 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] * 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org * 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s) * 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment * 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed * 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] * 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s) * 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie * 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment * 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] * 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage * 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie * 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie * 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet * 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]] * 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:14 cwilliams@dns1006: END - running authdns-update * 06:12 cwilliams@dns1006: START - running authdns-update * 06:11 cwilliams@dns1006: END - running authdns-update * 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json * 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage * 06:09 cwilliams@dns1006: START - running authdns-update * 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json * 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json * 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]] * 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet * 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json * 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]] * 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie * 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]] * 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]] * 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s) * 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment * 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s) * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed * 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie * 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration * 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw * 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw * 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage * 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie * 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111` == 2026-07-01 == * 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election * 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply * 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply * 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply * 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm * 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie * 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage * 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 * 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" * 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage * 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox * 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003 * 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm * 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie * 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie * 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage * 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie * 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage * 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie * 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" * 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie * 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage * 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage * 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage * 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie * 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie * 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) * 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment * 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] * 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie * 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts * 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts * 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet * 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet * 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt * 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet * 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet * 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie * 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]] * 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts * 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts * 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie * 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts * 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s) * 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage * 16:18 jasmine@dns1004: END - running authdns-update * 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART * 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage * 16:15 jasmine@dns1004: START - running authdns-update * 16:14 jasmine@dns1004: END - running authdns-update * 16:12 jasmine@dns1004: START - running authdns-update * 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance * 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde * 16:00 papaul: ongoing maintenance on lsw1-b2-codfw * 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie * 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt * 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie * 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet * 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover * 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . * 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie * 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply * 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply * 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply * 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply * 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply * 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]] * 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]] * 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply * 15:12 _joe_: restarted manually alertmanager-irc-relay * 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply * 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde * 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet * 15:02 papaul: ongoing maintenance on lsw1-a8-codfw * 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]] * 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) * 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment * 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] * 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover * 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]] * 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover * 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) * 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover * 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json * 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply * 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply * 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json * 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply * 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply * 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment * 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]] * 14:03 jmm@dns1004: END - running authdns-update * 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply * 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]] * 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad * 14:01 jmm@dns1004: START - running authdns-update * 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json * 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] * 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]] * 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]] * 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org * 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie * 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org * 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie * 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org * 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org * 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad * 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad * 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) * 13:30 caro@deploy1003: caro: Continuing with deployment * 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:27 topranks: route-engine failover cr1-eqiad * 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] * 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]] * 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) * 13:13 moritzm: installing qemu security updates * 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment * 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance * 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover * 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] * 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover * 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json * 13:01 moritzm: installing python3.13 security updates * 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json * 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]] * 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install * 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json * 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]] * 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet * 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet * 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie * 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie * 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) * 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed * 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage * 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage * 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]] * 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie * 11:50 cmooney@dns2005: END - running authdns-update * 11:49 cmooney@dns2005: START - running authdns-update * 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie * 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply * 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply * 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply * 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed * 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply * 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply * 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply * 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0) * 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" * 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply * 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply * 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie * 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie * 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply * 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply * 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox * 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie * 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage * 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie * 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage * 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage * 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet * 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]] * 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade * 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage * 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json * 10:26 moritzm: installing nginx security updates * 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie * 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json * 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]] * 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie * 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie * 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json * 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]] * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 * 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" * 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . * 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) * 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]] * 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org * 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) * 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment * 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org * 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . * 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] * 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) * 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] * 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) * 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) * 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] * 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) * 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment * 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] * 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) * 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] * 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] * 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) * 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment * 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. * 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] * 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]] * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 * 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 * 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" * 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie * 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues * 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie * 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 * 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 * 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage * 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie * 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage * 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie * 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie * 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage * 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie * 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash * 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie * 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie * 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage * 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie * 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage * 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage * 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie * 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) * 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie * 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image * 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json * 01:05 ladsgroup@dns1004: END - running authdns-update * 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json * 01:03 ladsgroup@dns1004: START - running authdns-update * 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json * 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]] * 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json * 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]] * 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json * 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie * 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie * 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) == Other archives == See [[Server Admin Log/Archives]]. <noinclude> [[Category:SAL]] [[Category:Operations]] </noinclude> 0l2w193a0xz501foeurueoinp7248wg Nova Resource:Admin/SAL 498 30942 2445260 2445061 2026-08-09T19:25:21Z Stashbot 7414 andrewbogott: "ceph osd out 224" in response to "osd.224 observed stalled read indications in DB device" 2445260 wikitext text/x-wiki === 2026-08-09 === * 19:25 andrewbogott: "ceph osd out 224" in response to "osd.224 observed stalled read indications in DB device" === 2026-08-06 === * 20:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_restart_mon_daemons (exit_code=0) * 20:24 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_restart_mon_daemons * 19:21 andrewbogott: restarting osds one by one to pick up config changes === 2026-08-05 === * 19:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,magnum * 19:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,magnum * 02:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) ([[phab:T429387|T429387]]) * 01:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T429387|T429387]]) * 01:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) ([[phab:T429387|T429387]]) * 01:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T429387|T429387]]) === 2026-08-04 === * 23:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) ([[phab:T429387|T429387]]) * 20:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T429387|T429387]]) * 20:03 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=97) ([[phab:T429387|T429387]]) * 20:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T429387|T429387]]) * 18:59 andrewbogott: systemctl reset-failed on cloudbackup2003. This patch should prevent future such false alarms: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1321058 === 2026-08-03 === * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) ([[phab:T431374|T431374]]) * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance ([[phab:T431374|T431374]]) * 14:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T431682|T431682]]) * 14:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T431682|T431682]]) * 14:50 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=97) ([[phab:T431374|T431374]]) * 14:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T431374|T431374]]) === 2026-07-28 === * 13:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) ([[phab:T431659|T431659]]) * 13:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 13:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons ([[phab:T431659|T431659]]) * 13:06 wmbot~dcaro@acme: END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) * 13:05 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 13:04 wmbot~dcaro@acme: END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) * 13:04 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 13:04 wmbot~dcaro@acme: END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) * 13:04 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 13:04 wmbot~dcaro@acme: END (ERROR) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=1) * 13:03 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 13:03 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 13:03 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 13:02 wmbot~dcaro@acme: END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) * 13:02 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 13:00 wmbot~dcaro@acme: END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) * 13:00 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 13:00 wmbot~dcaro@acme: END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) * 12:59 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:59 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 12:59 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:55 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 12:55 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:54 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 12:54 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:54 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 12:54 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:12 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 12:12 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:08 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 12:08 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:06 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 12:05 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 12:00 wmbot~dcaro@acme: END (ERROR) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=1) * 11:57 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 11:55 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 11:55 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 11:55 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) * 11:55 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.get_project_for_proxy * 03:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 02:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) * 02:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 02:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 01:57 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds === 2026-07-27 === * 22:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 18:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 18:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 18:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 16:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 15:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 15:51 andrewbogott: temporarily muting ceph slow ops alerts ("ceph health mute BLUESTORE_SLOW_OP_ALERT --sticky") for [[phab:T431659|T431659]] === 2026-07-24 === * 15:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) * 14:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds === 2026-07-23 === * 13:44 andrewbogott: restarting designate services in eqiad1; seeing many miscellaneous designate-sink errors in logs * 13:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,designate * 13:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate === 2026-07-22 === * 18:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T431374|T431374]]) * 18:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T431374|T431374]]) * 18:05 andrewbogott: serveraction powercycle on cloudvirt1071 * 18:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) ([[phab:T431374|T431374]]) * 18:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance ([[phab:T431374|T431374]]) * 12:46 taavi: add security group rules for all projects for new metricsinfra VIPs [[phab:T401813|T401813]] * 02:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 02:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2026-07-21 === * 15:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:22 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:38 godog: switch from SystemdUnitDown (eqiad only) to SystemdUnitFailed (codfw/eqiad) alerts -- there might be some codfw noise coming - [[phab:T428873|T428873]] === 2026-07-20 === * 21:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 21:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 21:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 21:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 21:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 21:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 21:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 20:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 20:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 20:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:07 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=99) * 12:07 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 12:01 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.roll_reboot_cloudnets (exit_code=0) * 11:25 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudservices.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::services<nowiki>}</nowiki>' * 11:19 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudservices.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::services<nowiki>}</nowiki>' * 10:13 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudlb.safe_reboot (exit_code=0) on hosts matched by 'A:cloudlb AND A:eqiad' * 10:08 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudlb.safe_reboot on hosts matched by 'A:cloudlb AND A:eqiad' * 10:02 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1011.eqiad.wmnet<nowiki>}</nowiki>' * 09:58 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1011.eqiad.wmnet<nowiki>}</nowiki>' * 09:53 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1007.eqiad.wmnet<nowiki>}</nowiki>' * 09:49 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1007.eqiad.wmnet<nowiki>}</nowiki>' * 09:26 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1006.eqiad.wmnet<nowiki>}</nowiki>' * 09:21 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1006.eqiad.wmnet<nowiki>}</nowiki>' === 2026-07-19 === * 16:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 16:31 andrewbogott: restarting eqiad1 openstack services; trying to resolve failures with allocating and deleting network interfaces * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2026-07-16 === * 14:17 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:17 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2026-07-15 === * 13:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T431429|T431429]]) * 13:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T431429|T431429]]) === 2026-07-14 === * 14:36 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::cloudweb<nowiki>}</nowiki>' * 14:31 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::cloudweb<nowiki>}</nowiki>' * 12:12 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1048' ([[phab:T431682|T431682]]) * 11:58 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1048' ([[phab:T431682|T431682]]) * 11:57 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) ([[phab:T431682|T431682]]) * 11:57 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance ([[phab:T431682|T431682]]) === 2026-07-09 === * 12:54 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/332 * 12:53 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/332 * 12:29 godog: wmcs-openstack aggregate delete maintenance * 10:11 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 10:11 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 10:11 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 10:10 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance === 2026-07-08 === * 19:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 19:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 19:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 19:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 17:21 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:21 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:20 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/326 * 17:20 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/326 * 17:05 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:04 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:03 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/323 * 17:03 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/323 * 16:58 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/322 * 16:58 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/322 * 16:57 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:56 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:55 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/321 * 16:55 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/321 * 16:53 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/321 * 16:53 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/321 * 16:49 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/321 * 16:49 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/321 * 13:52 andrewbogott: removing some cloudvirts from the 'maintenance' aggregate. I imagine they are there in error after some automated reboots. cloudvirt1049, cloudvirt1053, cloudvirt1064, cloudvirt1078, cloudvirt1080, cloudvirtlocal1001 === 2026-07-07 === * 09:14 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:13 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:07 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/319 * 09:06 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/319 * 04:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T431374|T431374]]) * 04:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T431374|T431374]]) * 04:18 andrewbogott: cycling power on cloudvirt1071 via mgmt/racadm; it seems unresponsive === 2026-07-06 === * 08:04 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.95<nowiki>}</nowiki>' * 07:08 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.95<nowiki>}</nowiki>' === 2026-06-29 === * 13:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 13:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2026-06-25 === * 06:51 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T424802|T424802]]) * 06:51 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T424802|T424802]]) === 2026-06-22 === * 18:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.nfs.migrate_service (exit_code=99) * 18:57 andrew@cloudcumin1001: START - Cookbook wmcs.nfs.migrate_service === 2026-06-17 === * 13:57 dhinus: updated wikireplicas-utils from 0.1.0 to 0.2.0 on clouddb* === 2026-06-16 === * 17:51 andrewbogott: ceph tell osd.126,127,129,131 compact * 17:04 andrewbogott: rebooting cloudcephosd1037 and cloudcephosd1038 because of missing volumes. The volumes reappeared after boot. * 16:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) ([[phab:T428385|T428385]]) * 16:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 16:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T429361|T429361]]) * 16:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T429361|T429361]]) * 16:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T429361|T429361]]) * 16:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T429361|T429361]]) * 16:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T429361|T429361]]) * 16:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T429361|T429361]]) * 16:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2010-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2010-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2011-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2011-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2006-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:33 andrewbogott: restarting quite a few other ceph-osd services in an attempt to quiet some slow op alerts * 15:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' ([[phab:T429361|T429361]]) * 15:13 andrewbogott: systemctl restart ceph-osd@58.service due to slow ops * 14:57 andrewbogott: "systemctl restart ceph-osd@281.service" as 281 shows as down * 14:55 andrewbogott: "systemctl restart ceph-osd@289.service" as 289 shows as down * 13:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudservices.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::services<nowiki>}</nowiki>' * 13:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudservices.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::services<nowiki>}</nowiki>' === 2026-06-15 === * 21:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 21:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 21:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 19:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 19:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 19:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 18:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 18:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 17:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 14:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 14:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_mons (exit_code=0) * 13:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_mons ([[phab:T428385|T428385]]) * 13:48 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.upgrade_osds (exit_code=97) * 13:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 13:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_mons (exit_code=99) * 12:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_mons ([[phab:T428385|T428385]]) === 2026-06-12 === * 14:00 volans: clearing cloudback-original snapshots of cinder volumes created in 2026 to allow the backups to not fail for [[phab:T428995|T428995]] === 2026-06-11 === * 17:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T428549|T428549]]) * 17:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T428549|T428549]]) * 17:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T428549|T428549]]) * 17:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T428549|T428549]]) * 16:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T428549|T428549]]) * 16:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T428549|T428549]]) * 16:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2006.eqiad.wmnet' ([[phab:T428549|T428549]]) * 16:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006.eqiad.wmnet' ([[phab:T428549|T428549]]) * 15:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2010-dev.codfw.wmnet' ([[phab:T428549|T428549]]) * 14:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2010-dev.codfw.wmnet' ([[phab:T428549|T428549]]) * 14:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2006-dev.codfw.wmnet' ([[phab:T428549|T428549]]) * 14:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' ([[phab:T428549|T428549]]) * 14:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T428549|T428549]]) * 14:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T428549|T428549]]) === 2026-06-09 === * 17:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 17:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T428385|T428385]]) * 16:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_mons (exit_code=0) * 16:35 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_mons ([[phab:T428385|T428385]]) === 2026-06-07 === * 17:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 17:40 andrewbogott: wmcs.openstack.restart_openstack --cluster-name eqiad1 --all as step one in troubleshooting [[phab:T428312|T428312]] * 17:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2026-06-02 === * 20:31 bd808: `kubectl sudo delete cm -n tool-arb-bot maintain-kubeusers-arb-bot` to trigger regeneration of .kube/config (IRC request) === 2026-05-26 === * 19:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) * 18:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 18:47 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=97) * 18:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 18:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) * 18:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 18:21 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=97) * 18:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 18:16 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=97) * 18:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 16:30 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' === 2026-05-25 === * 09:04 godog: move designate eqiad to zk backend === 2026-05-21 === * 18:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) * 18:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 18:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) * 18:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 17:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) * 17:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 14:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) * 14:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 14:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) * 14:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 14:17 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=97) * 14:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons * 13:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) ([[phab:T426563|T426563]]) * 13:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_mons ([[phab:T426563|T426563]]) * 11:59 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudlb.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudlb2002-dev.codfw.wmnet<nowiki>}</nowiki>' * 11:56 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudlb.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudlb2002-dev.codfw.wmnet<nowiki>}</nowiki>' === 2026-05-20 === * 13:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 12:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 12:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 12:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 12:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 11:56 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudlb.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudlb2002-dev.codfw.wmnet<nowiki>}</nowiki>' * 11:49 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudlb.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudlb2002-dev.codfw.wmnet<nowiki>}</nowiki>' * 11:28 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudlb.safe_reboot (exit_code=0) on hosts matched by 'A:cloudlb AND A:eqiad' * 11:16 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudlb.safe_reboot on hosts matched by 'A:cloudlb AND A:eqiad' * 11:15 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudlb.safe_reboot (exit_code=0) on hosts matched by 'A:cloudlb AND A:codfw' * 10:56 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudlb.safe_reboot on hosts matched by 'A:cloudlb AND A:codfw' * 10:53 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudlb.safe_reboot (exit_code=99) on hosts matched by 'A:cloudlb AND A:codfw' * 10:46 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudlb.safe_reboot on hosts matched by 'A:cloudlb AND A:codfw' * 10:46 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudlb.safe_reboot (exit_code=99) on hosts matched by 'A:cloudlb AND A:CODFW' * 10:46 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudlb.safe_reboot on hosts matched by 'A:cloudlb AND A:CODFW' * 01:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 01:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' === 2026-05-19 === * 22:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 22:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 22:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::control<nowiki>}</nowiki>' * 21:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::control<nowiki>}</nowiki>' * 18:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 18:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 17:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 17:58 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 17:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki> AND NOT P<nowiki>{</nowiki>F:kernelversion = 6.12.88<nowiki>}</nowiki>' * 11:57 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::cloudweb<nowiki>}</nowiki>' * 11:52 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::cloudweb<nowiki>}</nowiki>' * 11:51 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::cloudweb<nowiki>}</nowiki>' * 11:48 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::cloudweb<nowiki>}</nowiki>' * 11:13 volans@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 10:57 volans@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 10:43 volans@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 10:22 volans@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 10:21 volans@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 10:06 volans@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 09:55 volans@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 09:37 volans@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 02:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) ([[phab:T426563|T426563]]) * 02:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.reboot_node (exit_code=99) * 02:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.reboot_node * 00:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T426563|T426563]]) * 00:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) ([[phab:T426563|T426563]]) * 00:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T426563|T426563]]) === 2026-05-18 === * 23:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) ([[phab:T426563|T426563]]) * 20:35 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T426563|T426563]]) * 20:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) ([[phab:T426563|T426563]]) * 20:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T426563|T426563]]) * 20:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) ([[phab:T426563|T426563]]) * 20:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T426563|T426563]]) * 12:01 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/313 * 12:01 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/313 * 11:59 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:59 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:47 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/312 * 10:46 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/312 * 10:46 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/312 * 10:45 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/312 === 2026-05-11 === * 13:47 andrewbogott: restarting eqiad1 keystone services to pick up some logging changes in wmfkeystonehooks === 2026-05-06 === * 13:34 godog: change cloud-vps quota request phab at https://phabricator.wikimedia.org/project/manage/2880/ to mention https://cloudvps-quota.toolforge.org === 2026-05-05 === * 10:20 taavi: taavi@cloudcontrol1007 ~ $ sudo wmcs-enc-cli --openstack-project admin delete_project wikilabels # [[phab:T416588|T416588]] === 2026-05-02 === * 12:12 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:11 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2026-04-29 === * 16:05 andrewbogott: cleaning up stray broken osbpo references on VMs, e.g. /etc/apt/sources.list.d/openstack-dalmatian-bookworm.sources === 2026-04-28 === * 14:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,designate * 14:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 08:42 wmbot~godog@r5: END (PASS) - Cookbook wmcs.openstack.rack_resources (exit_code=0) on cluster 'codfw1dev' * 08:42 wmbot~godog@r5: START - Cookbook wmcs.openstack.rack_resources on cluster 'codfw1dev' * 08:32 wmbot~godog@r5: END (PASS) - Cookbook wmcs.openstack.rack_resources (exit_code=0) on cluster 'codfw1dev' * 08:32 wmbot~godog@r5: START - Cookbook wmcs.openstack.rack_resources on cluster 'codfw1dev' * 08:29 wmbot~godog@r5: END (PASS) - Cookbook wmcs.openstack.rack_resources (exit_code=0) on cluster 'codfw1dev' * 08:28 wmbot~godog@r5: START - Cookbook wmcs.openstack.rack_resources on cluster 'codfw1dev' * 08:19 wmbot~godog@r5: END (PASS) - Cookbook wmcs.openstack.rack_resources (exit_code=0) on cluster 'codfw1dev' * 08:19 wmbot~godog@r5: START - Cookbook wmcs.openstack.rack_resources on cluster 'codfw1dev' * 08:18 wmbot~godog@r5: END (PASS) - Cookbook wmcs.openstack.rack_resources (exit_code=0) on cluster 'codfw1dev' * 08:18 wmbot~godog@r5: START - Cookbook wmcs.openstack.rack_resources on cluster 'codfw1dev' === 2026-04-27 === * 12:43 godog: upgrade spicerack on cloudcumin === 2026-04-21 === * 15:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 15:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 15:30 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment codfw1dev for all services * 15:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2026-04-20 === * 10:20 godog: test shutting cloudcontrol2005-dev network port === 2026-04-16 === * 06:43 godog: roll-restart nova-api to pick up changes - [[phab:T423378|T423378]] === 2026-04-15 === * 15:02 godog: deploy per-service oslo.messaging shared memory file name - [[phab:T423378|T423378]] === 2026-04-14 === * 21:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 21:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 21:21 andrewbogott: rebuilding eqiad1 rabbitmq cluster in hopes of getting some more consistent api responses * 20:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova * 19:56 andrewbogott: restarting all nova services in eqiad1; i'm seeing inconsistent permission failures * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova * 12:45 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T419658|T419658]]) * 12:45 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T419658|T419658]]) * 12:30 godog: set maint on cloudvirt1050 - [[phab:T419658|T419658]] === 2026-04-13 === * 19:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) * 19:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 18:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 18:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 18:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 18:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 17:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 17:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 17:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 15:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 15:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 15:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 15:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 15:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 15:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) * 15:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds * 14:31 godog: grant filippo and volans admin roles in codfw1dev === 2026-04-08 === * 18:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 18:17 andrewbogott: restarting openstack services in eqiad1 before digging into a tofu failure * 18:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 14:03 godog: bounce nova on cloudcontrol1006 * 13:48 godog: stop designate and memcached on all cloudcontrol1* * 13:35 godog: leave designate processes up only on cloudcontrol1006 to ease debugging - [[phab:T422646|T422646]] * 11:59 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for service: project,designate * 11:58 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 11:54 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for service: project,designate * 11:53 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 09:45 godog: bounce designate on cloudcontrol * 08:14 godog: perform more network tests on cloudrabbit1001 - [[phab:T417393|T417393]] * 07:57 godog: unshut cloudcontrol1011 network interface - [[phab:T417393|T417393]] * 07:36 godog: shut cloudcontrol1011 network interface - [[phab:T417393|T417393]] * 07:26 godog: unshut cloudrabbit1001 network interface - [[phab:T417393|T417393]] * 07:00 godog: test shutting cloudrabbit1001 network interface - [[phab:T417393|T417393]] === 2026-04-07 === * 13:40 andrewbogott: upgrading spicerack on cloudcumin1001 === 2026-04-06 === * 13:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2026-04-03 === * 10:18 godog: move codfw neutron l3-agent queues to quorum - [[phab:T421054|T421054]] === 2026-04-02 === * 12:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 12:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 12:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:03 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:02 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:01 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/304 * 11:01 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/304 === 2026-04-01 === * 14:45 volans: Installed cumin v6.0.0-1 on apt.w.o (unattended upgrades) and the cloudcumin hosts (previous one for rollback in my home) * 12:24 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/304 * 12:23 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/304 * 08:15 godog: extend cloudrabbit1* root with an additional 100G - [[phab:T421054|T421054]] === 2026-03-31 === * 07:26 godog: neutron maint done - [[phab:T421054|T421054]] * 07:12 godog: start maint on neutron - [[phab:T421054|T421054]] === 2026-03-30 === * 13:52 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 13:48 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 13:45 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=0) on deployment codfw1dev * 13:42 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 11:53 godog: bounce neutron-l3-agent on cloudnet1005 * 11:20 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 11:06 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 11:01 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=0) on deployment eqiad1 * 10:57 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment eqiad1 * 10:36 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 10:21 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 09:58 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment eqiad1 * 09:57 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment eqiad1 * 09:42 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,designate * 09:41 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 09:30 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova,neutron,designate * 09:30 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for service: project,designate * 09:28 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 09:18 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova,neutron,designate * 09:18 filippo@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for service: project,neutron,designate * 09:15 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 09:14 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,neutron,designate * 09:14 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,heat * 09:13 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,heat * 09:12 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for all services * 09:11 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 09:11 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2026-03-26 === * 22:47 andrewbogott: I am logging to the admin log * 13:38 root@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 13:33 root@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2026-03-25 === * 17:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 17:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 17:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 17:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 17:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 17:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 16:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,heat * 16:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,heat === 2026-03-24 === * 10:39 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,nova,glance,keystone,cinder,neutron,trove,magnum,octavia,heat,swift * 10:35 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,nova,glance,keystone,cinder,neutron,trove,magnum,octavia,heat,swift * 10:34 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 10:32 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 10:31 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 10:30 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2026-03-23 === * 20:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirt1076.eqiad.wmnet' * 20:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1076.eqiad.wmnet' * 20:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirtlocal1076.eqiad.wmnet' * 20:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1076.eqiad.wmnet' * 20:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirtlocal1076.eqiad.wmnet' * 20:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1076.eqiad.wmnet' * 20:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirtlocal1076' * 20:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1076' * 16:24 dhinus: added komla to https://gitlab.wikimedia.org/groups/repos/cloud/-/group_members [[phab:T420532|T420532]] * 15:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 14:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 14:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 13:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 13:17 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,designate * 13:16 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 13:16 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,heat * 13:15 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,heat * 12:43 godog: apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1254877 to cloudrabbit eqiad - [[phab:T418444|T418444]] * 10:51 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 10:47 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 10:47 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project * 10:47 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project * 10:33 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=0) on deployment codfw1dev * 10:29 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 10:12 godog: apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1254877 to cloudrabbit codfw - [[phab:T418444|T418444]] === 2026-03-19 === * 16:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T419960|T419960]]) * 16:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) ([[phab:T419960|T419960]]) * 16:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.roll_reboot_osds ([[phab:T419960|T419960]]) === 2026-03-18 === * 20:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 20:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 18:43 andrewbogott: dist-upgrade and rebooting cloudrabbit2xxx-dev nodes * 09:26 godog: end network switch failover tests - [[phab:T417393|T417393]] * 08:58 godog: bounce rabbit on cloudrabbit1001 - [[phab:T417393|T417393]] * 08:55 filippo@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 08:51 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 08:00 godog: start network switch failover tests - [[phab:T417393|T417393]] === 2026-03-17 === * 13:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T406516|T406516]]) * 13:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T406516|T406516]]) * 13:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T406516|T406516]]) * 13:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T406516|T406516]]) * 13:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T406516|T406516]]) * 12:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T406516|T406516]]) * 01:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T406516|T406516]]) * 01:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T406516|T406516]]) * 01:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T406516|T406516]]) * 01:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T406516|T406516]]) * 01:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T406516|T406516]]) * 01:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T406516|T406516]]) * 01:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T406516|T406516]]) * 00:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T406516|T406516]]) === 2026-03-16 === * 23:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T406516|T406516]]) * 23:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T406516|T406516]]) * 23:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T406516|T406516]]) * 22:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1069.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1069.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1068.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1068.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T406516|T406516]]) * 21:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1070.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1070.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T406516|T406516]]) * 20:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1072.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1072.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1073.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1073.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1074.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1074.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 19:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 19:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1075.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1075.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvir1075.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvir1075.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudweb.unset_maintenance (exit_code=99) ([[phab:T406516|T406516]]) * 19:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.unset_maintenance ([[phab:T406516|T406516]]) * 19:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1005.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1005.eqiad.wmnet' ([[phab:T406516|T406516]]) * 19:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1006.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1006.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T406516|T406516]]) * 18:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T406516|T406516]]) * 17:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T406516|T406516]]) * 17:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=99) ([[phab:T406516|T406516]]) * 17:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T406516|T406516]]) * 17:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T406516|T406516]]) * 17:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T406516|T406516]]) * 17:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T406516|T406516]]) * 16:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T406516|T406516]]) * 05:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 05:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 05:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 04:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 04:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 04:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 04:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 04:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 04:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 03:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 03:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1045.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 03:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1045.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 03:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova * 01:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova * 01:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 01:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 01:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) === 2026-03-15 === * 03:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 03:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 03:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) === 2026-03-14 === * 02:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 02:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 01:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) === 2026-03-13 === * 23:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 23:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 23:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 23:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 22:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 22:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 22:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 22:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 22:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 22:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 21:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 21:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 21:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 21:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 21:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 21:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 20:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 20:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 20:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 20:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 20:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 20:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 19:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 19:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1068.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 19:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1068.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1069.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1069.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1070.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1070.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1071.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1071.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 17:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1072.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 16:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1073.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1074.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1074.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 16:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 15:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 15:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1076.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) * 15:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1076.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T419948|T419948]]) === 2026-03-12 === * 13:46 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:45 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:44 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:43 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:44 godog: create g4.cores8.ram32.disk20.ephem140 flavor for tools worker === 2026-03-10 === * 17:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 17:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2026-03-04 === * 14:43 taavi: deploying firewall rule updates: https://gerrit.wikimedia.org/r/c/operations/homer/public/+/970275 === 2026-02-26 === * 03:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 03:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 03:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=0) on deployment eqiad1 * 03:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment eqiad1 * 03:27 andrewbogott: rebuilding the rabbitmq cluster in eqiad1; many failed messages === 2026-02-24 === * 08:18 godog: clean up stray wikilabels project puppet host certs * 00:16 bd808: Applied LDIF to create analytics-sre user and group ([[phab:T418120|T418120]]) === 2026-02-20 === * 12:00 dhinus: DROP DATABASE toollabs_p; (was used by updatetools.py, see [[phab:T415383|T415383]]) === 2026-02-18 === * 13:29 taavi: rebooting cloudgw1004 for [[phab:T417075|T417075]] fixes === 2026-02-17 === * 12:28 volans: re-enabled puppet on toolforge's nfs k8s workers * 10:39 volans: temporarily disabling puppet on toolforge's nfs k8s workers to test gerrit/1239689 === 2026-02-13 === * 04:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 04:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 04:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=0) on deployment codfw1dev * 04:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev === 2026-02-12 === * 14:34 andrewbogott: moritz is upgrading dnsmasq on eqiad1 cloudnets and cloudvirts * 10:08 dhinus: delete job "updatetools" in admin tool as it's no longer used ([[phab:T415383|T415383]]) === 2026-02-09 === * 16:03 andrewbogott: rebooting cloudnets in codfw1dev to make sure we've picked up the new dnsmasq version * 12:50 dcaro: removed stall cert cloudinfra-acme-chief-01.novalocal from cloudinfra-internal-puppetserver-1 * 12:49 dcaro: removed the pki* certs from the cloudinfra-cloudvps puppetserver as they are not handled there anymore but the local puppetserver * 09:45 dcaro: re-enabling nrpe2nodexp-ferm_active.service on cloudcumins after upgrade (getting stall promfile) * 03:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 02:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 02:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=0) on deployment eqiad1 * 02:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment eqiad1 === 2026-02-05 === * 21:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 21:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=0) on deployment codfw1dev * 21:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 20:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment codfw1dev * 20:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 20:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment codfw1dev * 20:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 20:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment codfw1dev * 20:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 20:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment codfw1dev * 20:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 20:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment codfw1dev * 20:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 20:45 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment codfw1dev * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev * 20:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster (exit_code=99) on deployment codfw1dev * 20:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.rabbitmq.rebuild_rabbit_cluster on deployment codfw1dev === 2026-02-03 === * 18:02 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,designate * 18:01 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate === 2026-01-24 === * 17:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 17:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 17:09 andrewbogott: preemptively rebuilding rabbitmq for eqiad1; message flakiness === 2026-01-23 === * 13:02 taavi: switch https://gitlab.wikimedia.org/repos/cloud/wmcs/utils to fast-forward only merging mode === 2026-01-22 === * 18:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2006-dev.codfw.wmnet' * 18:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2006-dev.codfw.wmnet' * 18:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2005-dev.codfw.wmnet' * 18:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2005-dev.codfw.wmnet' * 17:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' * 17:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2004-dev.codfw.wmnet' * 17:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2005-dev.codfw.wmnet' * 17:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 17:03 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) on host 'cloudnet2006-dev.codfw.wmnet' * 17:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 15:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2006-dev.codfw.wmnet' * 15:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 14:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 14:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 14:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2004-dev.codfw.wmnet' * 14:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004-dev.codfw.wmnet' * 14:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2005-dev.codfw.wmnet' * 14:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 02:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2010-dev.codfw.wmnet' * 02:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2010-dev.codfw.wmnet' * 01:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2010-dev.codfw.wmnet' * 01:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2010-dev.codfw.wmnet' * 00:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2010-dev.codfw.wmnet' === 2026-01-21 === * 23:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2010-dev.codfw.wmnet' * 23:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2006-dev.codfw.wmnet' * 23:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' * 23:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2006-dev.codfw.wmnet' * 23:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' * 23:18 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) on host 'cloudcontrol2006-dev.codfw.wmnet' (Txxxxxx) * 23:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' (Txxxxxx) * 23:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2006-dev.codfw.wmnet' (Txxxxxx) * 22:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' (Txxxxxx) * 22:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' (Txxxxxx) * 22:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' (Txxxxxx) === 2026-01-14 === * 15:43 andrewbogott: reimaging cloudgw2003-dev to Trixie * 13:37 andrewbogott: reimaging cloudgw2002-dev to Trixie === 2026-01-12 === * 18:09 andrewbogott: stopping bird on cloudlb1001, reimaging to Trixie * 15:54 andrewbogott: stopping bird on cloudlb1002, reimaging to Trixie * 13:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 13:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2026-01-10 === * 20:56 andrewbogott: in codfw1dev: openstack role add --project swift --user swift member * 20:56 andrewbogott: in codfw1dev: openstack role add --project swift --user swift admin * 19:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 19:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2026-01-07 === * 10:00 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 10:00 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2026-01-06 === * 02:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 02:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2026-01-05 === * 22:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova * 22:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova === 2025-12-26 === * 23:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 23:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-12-20 === * 03:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 03:01 andrewbogott: restarting eqiad1 openstack services in response to some fullstack failures * 02:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 02:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,heat * 02:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,heat === 2025-12-19 === * 16:18 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch ([[phab:T412865|T412865]]) * 16:17 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch ([[phab:T412865|T412865]]) === 2025-12-18 === * 11:04 godog: bump tools object quota to 500G === 2025-12-16 === * 18:51 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 18:51 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 18:50 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/286 * 18:50 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/286 * 18:48 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/286 * 18:47 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/286 * 13:18 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.reboot_node (exit_code=0) * 13:03 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.reboot_node === 2025-12-05 === * 15:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 15:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-12-03 === * 10:55 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:55 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:48 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/285 * 10:48 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/285 * 05:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 05:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 00:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 00:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2025-12-02 === * 23:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 23:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 23:04 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 23:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 23:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for all services * 23:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 20:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1061.eqiad.wmnet' * 20:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 1797e3a3-f04a-4cb0-9102-{{Gerrit|79f1d0079d57}} (cluster eqiad1) * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 1797e3a3-f04a-4cb0-9102-{{Gerrit|79f1d0079d57}} (cluster eqiad1) * 20:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 592754fe-dd64-463f-ab33-{{Gerrit|d51a4108cec0}} (cluster eqiad1) * 20:50 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 592754fe-dd64-463f-ab33-{{Gerrit|d51a4108cec0}} (cluster eqiad1) * 20:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 20:32 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 20:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 20:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' * 20:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' * 20:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm c6645048-8447-4553-bf13-{{Gerrit|8122f959e4a8}} (cluster eqiad1) * 20:23 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm c6645048-8447-4553-bf13-{{Gerrit|8122f959e4a8}} (cluster eqiad1) * 20:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1054.eqiad.wmnet' * 20:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' * 20:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1054.eqiad.wmnet' * 20:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' * 20:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' * 20:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1050.eqiad.wmnet' * 20:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' * 20:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 20:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1046.eqiad.wmnet' * 20:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 19:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1046.eqiad.wmnet' * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 19:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1046.eqiad.wmnet' * 19:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 17:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 17:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 16:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 16:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 1da6f8f7-db35-4f33-92f9-{{Gerrit|29a6516bf47c}} (cluster eqiad1) * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 1da6f8f7-db35-4f33-92f9-{{Gerrit|29a6516bf47c}} (cluster eqiad1) * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 3f0dc3e0-f5e8-43a4-86dc-{{Gerrit|523ad08e90e6}} (cluster eqiad1) * 16:51 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 3f0dc3e0-f5e8-43a4-86dc-{{Gerrit|523ad08e90e6}} (cluster eqiad1) * 16:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 848de230-1687-40c5-b954-{{Gerrit|f8c2a3b7a443}} (cluster eqiad1) * 16:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 16:50 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 848de230-1687-40c5-b954-{{Gerrit|f8c2a3b7a443}} (cluster eqiad1) * 16:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 9491619f-43b5-4612-b976-{{Gerrit|00862dcd901d}} (cluster eqiad1) * 16:49 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 9491619f-43b5-4612-b976-{{Gerrit|00862dcd901d}} (cluster eqiad1) * 16:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 59f99bab-8a86-4701-a142-{{Gerrit|3a15a1c18d48}} (cluster eqiad1) * 16:48 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 59f99bab-8a86-4701-a142-{{Gerrit|3a15a1c18d48}} (cluster eqiad1) * 16:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm afb538cb-a128-450b-a02f-{{Gerrit|4fee25183588}} (cluster eqiad1) * 16:47 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm afb538cb-a128-450b-a02f-{{Gerrit|4fee25183588}} (cluster eqiad1) * 16:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 8b546fd2-137d-4b91-86f3-{{Gerrit|b50fa515c98c}} (cluster eqiad1) * 16:46 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 8b546fd2-137d-4b91-86f3-{{Gerrit|b50fa515c98c}} (cluster eqiad1) * 16:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm c7b1311c-ee8b-4118-b907-{{Gerrit|ad0382644350}} (cluster eqiad1) * 16:46 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm c7b1311c-ee8b-4118-b907-{{Gerrit|ad0382644350}} (cluster eqiad1) * 16:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 60d1c96e-0c3c-47e1-86d6-{{Gerrit|cd30527d5066}} (cluster eqiad1) * 16:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 16:45 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 60d1c96e-0c3c-47e1-86d6-{{Gerrit|cd30527d5066}} (cluster eqiad1) * 16:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 028c29da-adcb-4239-bcb4-{{Gerrit|6e80516e6fbb}} (cluster eqiad1) * 16:44 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 028c29da-adcb-4239-bcb4-{{Gerrit|6e80516e6fbb}} (cluster eqiad1) * 16:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 16:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm e6796dbd-2511-4bf6-bdee-{{Gerrit|4a14a7414d5f}} (cluster eqiad1) * 16:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 5e5c3bad-f1c7-49e5-b846-{{Gerrit|edaf111af83c}} (cluster eqiad1) * 16:33 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm e6796dbd-2511-4bf6-bdee-{{Gerrit|4a14a7414d5f}} (cluster eqiad1) * 16:33 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 5e5c3bad-f1c7-49e5-b846-{{Gerrit|edaf111af83c}} (cluster eqiad1) * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 1d74fc9a-0ddd-41d6-a0fd-{{Gerrit|5bba5e455c32}} (cluster eqiad1) * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 55ed5d49-43db-4f62-8c40-{{Gerrit|5cb0431dfce2}} (cluster eqiad1) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 1d74fc9a-0ddd-41d6-a0fd-{{Gerrit|5bba5e455c32}} (cluster eqiad1) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 55ed5d49-43db-4f62-8c40-{{Gerrit|5cb0431dfce2}} (cluster eqiad1) * 15:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 8030caca-e1e8-4f1d-bce1-{{Gerrit|04afd22adb3a}} (cluster eqiad1) * 15:48 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 8030caca-e1e8-4f1d-bce1-{{Gerrit|04afd22adb3a}} (cluster eqiad1) * 15:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 9ccf684d-c6ea-45ee-83db-{{Gerrit|ee3af5de3dfe}} (cluster eqiad1) * 15:47 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 9ccf684d-c6ea-45ee-83db-{{Gerrit|ee3af5de3dfe}} (cluster eqiad1) * 15:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 02bf16d5-5e10-470b-b05d-{{Gerrit|341673a284de}} (cluster eqiad1) * 15:47 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 02bf16d5-5e10-470b-b05d-{{Gerrit|341673a284de}} (cluster eqiad1) * 15:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 106a0f58-3276-4754-93cd-{{Gerrit|a7ae20fddc75}} (cluster eqiad1) * 15:46 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 106a0f58-3276-4754-93cd-{{Gerrit|a7ae20fddc75}} (cluster eqiad1) * 15:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm d9e4c884-82f1-4c2e-8b35-{{Gerrit|70bfeb5292cf}} (cluster eqiad1) * 15:46 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm d9e4c884-82f1-4c2e-8b35-{{Gerrit|70bfeb5292cf}} (cluster eqiad1) * 15:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 4945aa99-aeff-4198-9aaa-{{Gerrit|7391c9a84c55}} (cluster eqiad1) * 15:45 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 4945aa99-aeff-4198-9aaa-{{Gerrit|7391c9a84c55}} (cluster eqiad1) * 15:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 278b5002-e9db-4506-a40f-{{Gerrit|167b52b9515f}} (cluster eqiad1) * 15:45 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 278b5002-e9db-4506-a40f-{{Gerrit|167b52b9515f}} (cluster eqiad1) * 15:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 36b6590c-eac2-40a0-ac30-{{Gerrit|7cf79ff12ce3}} (cluster eqiad1) * 15:44 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 36b6590c-eac2-40a0-ac30-{{Gerrit|7cf79ff12ce3}} (cluster eqiad1) * 15:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 797c51db-bc81-4363-922e-{{Gerrit|a52c3fc3eeea}} (cluster eqiad1) * 15:43 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 797c51db-bc81-4363-922e-{{Gerrit|a52c3fc3eeea}} (cluster eqiad1) * 15:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 627735e5-57c7-4714-855b-{{Gerrit|b7311fc527c6}} (cluster eqiad1) * 15:43 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 627735e5-57c7-4714-855b-{{Gerrit|b7311fc527c6}} (cluster eqiad1) * 15:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 90b1b0c3-14fb-47f6-9c50-{{Gerrit|f952f55bcfea}} (cluster eqiad1) * 15:42 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 90b1b0c3-14fb-47f6-9c50-{{Gerrit|f952f55bcfea}} (cluster eqiad1) * 15:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm bc10309b-5227-4ca2-b74c-{{Gerrit|440e2fdc116e}} (cluster eqiad1) * 15:40 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm bc10309b-5227-4ca2-b74c-{{Gerrit|440e2fdc116e}} (cluster eqiad1) * 15:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm c4c0ffc0-ebd2-4133-9912-{{Gerrit|585af2725bfd}} (cluster eqiad1) * 15:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 2f7b7cfa-12ed-41d5-977d-{{Gerrit|1e11e8335cf4}} (cluster eqiad1) * 15:40 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm c4c0ffc0-ebd2-4133-9912-{{Gerrit|585af2725bfd}} (cluster eqiad1) * 15:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 82b22752-6752-4814-90c5-{{Gerrit|2aebd3825e95}} (cluster eqiad1) * 15:39 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 2f7b7cfa-12ed-41d5-977d-{{Gerrit|1e11e8335cf4}} (cluster eqiad1) * 15:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm d761eca2-4a21-4522-8d95-{{Gerrit|584bf639e6c0}} (cluster eqiad1) * 15:39 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 82b22752-6752-4814-90c5-{{Gerrit|2aebd3825e95}} (cluster eqiad1) * 15:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 6c263c50-71da-40ee-b1e0-{{Gerrit|00d40ba108e7}} (cluster eqiad1) * 15:38 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 6c263c50-71da-40ee-b1e0-{{Gerrit|00d40ba108e7}} (cluster eqiad1) * 15:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 65283e58-53e0-4545-b201-{{Gerrit|dab88a8ae7e5}} (cluster eqiad1) * 15:38 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm d761eca2-4a21-4522-8d95-{{Gerrit|584bf639e6c0}} (cluster eqiad1) * 15:37 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 65283e58-53e0-4545-b201-{{Gerrit|dab88a8ae7e5}} (cluster eqiad1) * 15:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm b875763a-d70f-4cab-92ce-{{Gerrit|60a523161799}} (cluster eqiad1) * 15:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm abf4e1e6-1bd6-41f2-ad1c-{{Gerrit|345e940b0b8b}} (cluster eqiad1) * 15:36 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm abf4e1e6-1bd6-41f2-ad1c-{{Gerrit|345e940b0b8b}} (cluster eqiad1) * 15:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 43462c80-0923-4494-a5db-{{Gerrit|a8df39d71cdd}} (cluster eqiad1) * 15:35 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 43462c80-0923-4494-a5db-{{Gerrit|a8df39d71cdd}} (cluster eqiad1) * 15:35 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm b875763a-d70f-4cab-92ce-{{Gerrit|60a523161799}} (cluster eqiad1) * 15:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm a80c58d9-fcce-4739-9f83-{{Gerrit|204cff354959}} (cluster eqiad1) * 15:34 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm a80c58d9-fcce-4739-9f83-{{Gerrit|204cff354959}} (cluster eqiad1) * 15:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 03d3daa4-c46e-4152-a4dd-{{Gerrit|c02a872f7edd}} (cluster eqiad1) * 15:34 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 03d3daa4-c46e-4152-a4dd-{{Gerrit|c02a872f7edd}} (cluster eqiad1) * 15:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 2074763a-97af-4b3d-a3b5-{{Gerrit|7d5cf43b9ecd}} (cluster eqiad1) * 15:33 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 2074763a-97af-4b3d-a3b5-{{Gerrit|7d5cf43b9ecd}} (cluster eqiad1) * 15:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 23c93ff7-f301-41e5-9ea5-{{Gerrit|9d4b2da1bf22}} (cluster eqiad1) * 15:33 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 23c93ff7-f301-41e5-9ea5-{{Gerrit|9d4b2da1bf22}} (cluster eqiad1) * 15:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 24e5b10a-80df-4bbc-807c-{{Gerrit|97d4e935d1f4}} (cluster eqiad1) * 15:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 35433ec4-9fd5-49f8-ac51-{{Gerrit|c05ecb433a4d}} (cluster eqiad1) * 15:32 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 24e5b10a-80df-4bbc-807c-{{Gerrit|97d4e935d1f4}} (cluster eqiad1) * 15:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 8e6f5b87-57b8-4ba5-b9e6-{{Gerrit|8feb4e413f3d}} (cluster eqiad1) * 15:32 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 35433ec4-9fd5-49f8-ac51-{{Gerrit|c05ecb433a4d}} (cluster eqiad1) * 15:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 01042fb6-b2e4-4690-88fb-{{Gerrit|3840c98b01aa}} (cluster eqiad1) * 15:31 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 8e6f5b87-57b8-4ba5-b9e6-{{Gerrit|8feb4e413f3d}} (cluster eqiad1) * 15:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm be716e27-6b34-4cb0-a498-{{Gerrit|b300937edc4c}} (cluster eqiad1) * 15:31 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 01042fb6-b2e4-4690-88fb-{{Gerrit|3840c98b01aa}} (cluster eqiad1) * 15:30 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm be716e27-6b34-4cb0-a498-{{Gerrit|b300937edc4c}} (cluster eqiad1) * 15:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 736c7c6d-319d-43e0-b2b1-{{Gerrit|efdd84b4736a}} (cluster eqiad1) * 15:30 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 736c7c6d-319d-43e0-b2b1-{{Gerrit|efdd84b4736a}} (cluster eqiad1) * 15:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm a5cb6818-f3ac-4ba9-afb5-{{Gerrit|5c657cf65f9a}} (cluster eqiad1) * 15:29 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm a5cb6818-f3ac-4ba9-afb5-{{Gerrit|5c657cf65f9a}} (cluster eqiad1) * 15:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 702a55e1-e176-45f1-af81-{{Gerrit|569013f91be3}} (cluster eqiad1) * 15:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm a5705913-72f2-4abd-84e6-{{Gerrit|3e084bfbd98d}} (cluster eqiad1) * 15:29 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 702a55e1-e176-45f1-af81-{{Gerrit|569013f91be3}} (cluster eqiad1) * 15:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 2a4b9dfd-7006-4b5b-8c95-{{Gerrit|7883709e5b2d}} (cluster eqiad1) * 15:28 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 2a4b9dfd-7006-4b5b-8c95-{{Gerrit|7883709e5b2d}} (cluster eqiad1) * 15:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 5037f71d-bcbf-4ed7-809b-{{Gerrit|052ca6026219}} (cluster eqiad1) * 15:28 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 5037f71d-bcbf-4ed7-809b-{{Gerrit|052ca6026219}} (cluster eqiad1) * 15:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 31bf05f8-122b-4558-8932-{{Gerrit|7ac4b8375ed5}} (cluster eqiad1) * 15:27 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 31bf05f8-122b-4558-8932-{{Gerrit|7ac4b8375ed5}} (cluster eqiad1) * 15:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 0b5b7c51-dc42-4bec-90f2-{{Gerrit|161807a385f7}} (cluster eqiad1) * 15:27 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm a5705913-72f2-4abd-84e6-{{Gerrit|3e084bfbd98d}} (cluster eqiad1) * 15:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 966b6ab3-b561-4e46-bfcd-{{Gerrit|1681ce9e91ac}} (cluster eqiad1) * 15:26 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 0b5b7c51-dc42-4bec-90f2-{{Gerrit|161807a385f7}} (cluster eqiad1) * 15:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm f7e8f001-e9c0-4fe9-8887-{{Gerrit|32289702b804}} (cluster eqiad1) * 15:26 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm f7e8f001-e9c0-4fe9-8887-{{Gerrit|32289702b804}} (cluster eqiad1) * 15:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm d73171e3-49ef-4d40-8008-{{Gerrit|a900781ea102}} (cluster eqiad1) * 15:26 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 966b6ab3-b561-4e46-bfcd-{{Gerrit|1681ce9e91ac}} (cluster eqiad1) * 15:25 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm d73171e3-49ef-4d40-8008-{{Gerrit|a900781ea102}} (cluster eqiad1) * 15:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 63a229be-765f-4b48-b8d9-{{Gerrit|24ee39243604}} (cluster eqiad1) * 15:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 9a7cf939-c634-4aa1-9fd2-{{Gerrit|dbc14b18d70e}} (cluster eqiad1) * 15:25 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 63a229be-765f-4b48-b8d9-{{Gerrit|24ee39243604}} (cluster eqiad1) * 15:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 63f68215-d302-4684-a91e-{{Gerrit|58f5272486a5}} (cluster eqiad1) * 15:25 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 9a7cf939-c634-4aa1-9fd2-{{Gerrit|dbc14b18d70e}} (cluster eqiad1) * 15:24 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 63f68215-d302-4684-a91e-{{Gerrit|58f5272486a5}} (cluster eqiad1) * 15:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 9c66be4e-6787-4844-a9ff-{{Gerrit|a65295ac5aac}} (cluster eqiad1) * 15:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm a4892d89-0981-412e-9f00-{{Gerrit|8882416948a1}} (cluster eqiad1) * 15:23 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm a4892d89-0981-412e-9f00-{{Gerrit|8882416948a1}} (cluster eqiad1) * 15:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm f8e08f70-f87e-413d-acad-{{Gerrit|080126ad5b1a}} (cluster eqiad1) * 15:23 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 9c66be4e-6787-4844-a9ff-{{Gerrit|a65295ac5aac}} (cluster eqiad1) * 15:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm e6796dbd-2511-4bf6-bdee-{{Gerrit|4a14a7414d5f}} (cluster eqiad1) * 15:23 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm f8e08f70-f87e-413d-acad-{{Gerrit|080126ad5b1a}} (cluster eqiad1) * 15:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 7e3011e8-aed8-4bed-8e18-{{Gerrit|f75afe3ec3a2}} (cluster eqiad1) * 15:22 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm e6796dbd-2511-4bf6-bdee-{{Gerrit|4a14a7414d5f}} (cluster eqiad1) * 15:22 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 7e3011e8-aed8-4bed-8e18-{{Gerrit|f75afe3ec3a2}} (cluster eqiad1) * 15:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 6fa9b0be-219d-4b10-962e-{{Gerrit|fa3a71f6740c}} (cluster eqiad1) * 15:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 5e5c3bad-f1c7-49e5-b846-{{Gerrit|edaf111af83c}} (cluster eqiad1) * 15:21 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 5e5c3bad-f1c7-49e5-b846-{{Gerrit|edaf111af83c}} (cluster eqiad1) * 15:21 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 6fa9b0be-219d-4b10-962e-{{Gerrit|fa3a71f6740c}} (cluster eqiad1) * 15:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 995817db-0966-485d-aca5-{{Gerrit|e5377c77a005}} (cluster eqiad1) * 15:21 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 995817db-0966-485d-aca5-{{Gerrit|e5377c77a005}} (cluster eqiad1) * 15:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm d409f39a-e24a-462e-b588-{{Gerrit|6f5f6557e26b}} (cluster eqiad1) * 15:20 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm d409f39a-e24a-462e-b588-{{Gerrit|6f5f6557e26b}} (cluster eqiad1) * 15:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 7dc8757e-8b8d-4cc9-ac8e-{{Gerrit|a2f925639f0b}} (cluster eqiad1) * 15:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.vps.instance.stop_start (exit_code=99) vm ci2.mediawiki-quickstart.eqiad1.wikimedia.cloud (cluster eqiad1) * 15:19 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm ci2.mediawiki-quickstart.eqiad1.wikimedia.cloud (cluster eqiad1) * 15:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.vps.instance.stop_start (exit_code=99) vm None (cluster eqiad1) * 15:19 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm None (cluster eqiad1) * 15:19 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 7dc8757e-8b8d-4cc9-ac8e-{{Gerrit|a2f925639f0b}} (cluster eqiad1) * 15:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 94f8be5c-3cdf-47cb-80b2-{{Gerrit|43c44da01789}} (cluster eqiad1) * 15:18 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 94f8be5c-3cdf-47cb-80b2-{{Gerrit|43c44da01789}} (cluster eqiad1) * 15:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm b89a2d14-2bc3-481b-baf6-{{Gerrit|496edadbe242}} (cluster eqiad1) * 15:17 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm b89a2d14-2bc3-481b-baf6-{{Gerrit|496edadbe242}} (cluster eqiad1) * 15:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm a8269e6b-e09d-4b6b-909e-{{Gerrit|5e4014165440}} (cluster eqiad1) * 15:17 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm a8269e6b-e09d-4b6b-909e-{{Gerrit|5e4014165440}} (cluster eqiad1) * 15:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm b4b2d7cd-3492-4d04-86ce-{{Gerrit|4c0b8344ddc3}} (cluster eqiad1) * 15:16 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm b4b2d7cd-3492-4d04-86ce-{{Gerrit|4c0b8344ddc3}} (cluster eqiad1) * 15:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm def4772e-c01f-4e5e-8e70-{{Gerrit|d62a546ebc2a}} (cluster eqiad1) * 15:16 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm def4772e-c01f-4e5e-8e70-{{Gerrit|d62a546ebc2a}} (cluster eqiad1) * 15:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 2b573025-a221-482a-b6b7-{{Gerrit|fd7e1b1308f7}} (cluster eqiad1) * 15:15 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 2b573025-a221-482a-b6b7-{{Gerrit|fd7e1b1308f7}} (cluster eqiad1) * 15:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 55b557f8-5817-47f6-adb4-{{Gerrit|abccac2b2997}} (cluster eqiad1) * 15:15 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 55b557f8-5817-47f6-adb4-{{Gerrit|abccac2b2997}} (cluster eqiad1) * 15:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 62cf19eb-ebf9-49d9-baa7-{{Gerrit|fba2cd6942d6}} (cluster eqiad1) * 15:14 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 62cf19eb-ebf9-49d9-baa7-{{Gerrit|fba2cd6942d6}} (cluster eqiad1) * 15:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm d3812466-439a-4355-901c-{{Gerrit|b1097a033d0b}} (cluster eqiad1) * 15:13 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm d3812466-439a-4355-901c-{{Gerrit|b1097a033d0b}} (cluster eqiad1) * 15:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 06ea2eca-b4e5-42fa-afc3-{{Gerrit|684b3c2b87a3}} (cluster eqiad1) * 15:13 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 06ea2eca-b4e5-42fa-afc3-{{Gerrit|684b3c2b87a3}} (cluster eqiad1) * 15:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 7cafb806-ecc1-459d-b6b3-{{Gerrit|4213511f1257}} (cluster eqiad1) * 15:12 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 7cafb806-ecc1-459d-b6b3-{{Gerrit|4213511f1257}} (cluster eqiad1) * 15:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm bbe5835d-5e8d-4778-b80f-{{Gerrit|5c6424928a88}} (cluster eqiad1) * 15:12 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm bbe5835d-5e8d-4778-b80f-{{Gerrit|5c6424928a88}} (cluster eqiad1) * 15:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 143f7189-22b2-4708-9a33-{{Gerrit|91d368a5c8eb}} (cluster eqiad1) * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 143f7189-22b2-4708-9a33-{{Gerrit|91d368a5c8eb}} (cluster eqiad1) * 15:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm b5fcdad2-d7d9-4ba2-903e-{{Gerrit|188231b05f71}} (cluster eqiad1) * 15:09 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm b5fcdad2-d7d9-4ba2-903e-{{Gerrit|188231b05f71}} (cluster eqiad1) * 15:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 9c7abfa9-a848-4ed5-8abb-{{Gerrit|dda2cb842b04}} (cluster eqiad1) * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 9c7abfa9-a848-4ed5-8abb-{{Gerrit|dda2cb842b04}} (cluster eqiad1) * 15:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm dedb5d73-e8f8-47d3-8598-{{Gerrit|2900be096236}} (cluster eqiad1) * 15:06 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm dedb5d73-e8f8-47d3-8598-{{Gerrit|2900be096236}} (cluster eqiad1) * 15:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 7e7b75d2-0f13-4973-ad18-{{Gerrit|3dc7a52b0781}} (cluster eqiad1) * 15:06 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 7e7b75d2-0f13-4973-ad18-{{Gerrit|3dc7a52b0781}} (cluster eqiad1) * 15:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 9b6c52a7-527d-4513-b996-{{Gerrit|160af646c5fb}} (cluster eqiad1) * 15:05 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 9b6c52a7-527d-4513-b996-{{Gerrit|160af646c5fb}} (cluster eqiad1) * 15:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm d4f8ac4d-3059-499a-961c-{{Gerrit|505f6ff89675}} (cluster eqiad1) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm d4f8ac4d-3059-499a-961c-{{Gerrit|505f6ff89675}} (cluster eqiad1) * 15:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 2a87ab18-417b-42a6-9ee1-{{Gerrit|a273a2379e62}} (cluster eqiad1) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 2a87ab18-417b-42a6-9ee1-{{Gerrit|a273a2379e62}} (cluster eqiad1) * 15:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 9f04369c-d074-44e2-a3b2-{{Gerrit|b0545accd0e0}} (cluster eqiad1) * 15:03 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 9f04369c-d074-44e2-a3b2-{{Gerrit|b0545accd0e0}} (cluster eqiad1) * 15:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 1b6df367-1d3d-4e48-8333-{{Gerrit|7a4f79a49a2a}} (cluster eqiad1) * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 1b6df367-1d3d-4e48-8333-{{Gerrit|7a4f79a49a2a}} (cluster eqiad1) * 15:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm a07ac179-366e-49a4-9499-{{Gerrit|bee721949963}} (cluster eqiad1) * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm a07ac179-366e-49a4-9499-{{Gerrit|bee721949963}} (cluster eqiad1) === 2025-11-27 === * 04:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 04:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 04:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 04:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 04:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 04:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 04:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 04:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-11-26 === * 15:26 dhinus: depool clouddb10[17-20] for network maintenance [[phab:T404609|T404609]] * 10:59 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 10:59 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 10:24 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch ([[phab:T408387|T408387]]) * 10:23 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch ([[phab:T408387|T408387]]) * 10:23 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 10:23 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2025-11-25 === * 15:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 15:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 15:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 15:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 15:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 15:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 15:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 15:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 08:17 godog: restart neutron-metadata-agent for testing - [[phab:T410983|T410983]] === 2025-11-24 === * 15:01 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:01 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 14:53 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 14:53 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 14:30 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 14:29 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2025-11-23 === * 21:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 20:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 20:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 20:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 20:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2005-dev.codfw.wmnet' * 19:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2005-dev.codfw.wmnet' === 2025-11-19 === * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' * 20:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 20:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2004-dev.codfw.wmnet' * 20:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2004-dev.codfw.wmnet' * 20:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 20:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 20:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2004-dev.codfw.wmnet' * 19:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1070.eqiad.wmnet' * 19:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1070.eqiad.wmnet' * 19:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.instance.stop_start (exit_code=0) vm 764d92cd-09df-468a-9595-{{Gerrit|7bcfbc4a8841}} (cluster eqiad1) * 19:44 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm 764d92cd-09df-468a-9595-{{Gerrit|7bcfbc4a8841}} (cluster eqiad1) * 19:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.vps.instance.stop_start (exit_code=99) vm content-diff-index (cluster eqiad1) * 19:44 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm content-diff-index (cluster eqiad1) * 19:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.vps.instance.stop_start (exit_code=99) vm None (cluster eqiad1) * 19:44 andrew@cloudcumin1001: START - Cookbook wmcs.vps.instance.stop_start vm None (cluster eqiad1) * 18:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1040.eqiad.wmnet' * 18:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' * 18:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1040.eqiad.wmnet' * 18:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' * 18:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1040.eqiad.wmnet' * 18:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' * 18:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1040.eqiad.wmnet' * 18:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' * 16:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2004-dev.codfw.wmnet' * 16:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2004-dev.codfw.wmnet' * 16:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1070.eqiad.wmnet' * 16:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1070.eqiad.wmnet' * 16:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 15:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 15:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1054.eqiad.wmnet' * 15:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' * 15:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1054.eqiad.wmnet' * 15:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' * 15:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1046.eqiad.wmnet' * 15:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 15:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1050.eqiad.wmnet' * 15:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1050.eqiad.wmnet' * 15:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 15:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 15:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 15:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 15:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 15:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1050.eqiad.wmnet' * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1050.eqiad.wmnet' * 03:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1050.eqiad.wmnet' * 03:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1050.eqiad.wmnet' * 03:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1046.eqiad.wmnet' * 03:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 01:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1076.eqiad.wmnet' * 01:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1076.eqiad.wmnet' * 01:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1046.eqiad.wmnet' * 01:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' === 2025-11-18 === * 22:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1076.eqiad.wmnet' * 22:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1076.eqiad.wmnet' * 22:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1076.eqiad.wmnet' * 21:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1076.eqiad.wmnet' * 21:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1075.eqiad.wmnet' * 21:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1075.eqiad.wmnet' * 21:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1074.eqiad.wmnet' * 21:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1074.eqiad.wmnet' * 21:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1073.eqiad.wmnet' * 21:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1073.eqiad.wmnet' * 21:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1073.eqiad.wmnet' * 21:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1073.eqiad.wmnet' * 21:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1073.eqiad.wmnet' * 20:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1073.eqiad.wmnet' * 20:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1072.eqiad.wmnet' * 19:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1072.eqiad.wmnet' * 19:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1071.eqiad.wmnet' * 19:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1071.eqiad.wmnet' * 19:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1070.eqiad.wmnet' * 19:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1070.eqiad.wmnet' * 19:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1070.eqiad.wmnet' * 19:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1070.eqiad.wmnet' * 19:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1069.eqiad.wmnet' * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1069.eqiad.wmnet' * 18:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1068.eqiad.wmnet' * 18:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1068.eqiad.wmnet' * 18:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1067.eqiad.wmnet' * 18:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1067.eqiad.wmnet' * 18:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1066.eqiad.wmnet' * 17:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1066.eqiad.wmnet' * 17:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' * 17:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 17:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1044.eqiad.wmnet' * 17:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 17:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1044.eqiad.wmnet' * 17:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 17:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1065.eqiad.wmnet' * 17:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1065.eqiad.wmnet' * 17:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1064.eqiad.wmnet' * 16:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1064.eqiad.wmnet' * 16:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1063.eqiad.wmnet' * 16:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1063.eqiad.wmnet' * 16:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1062.eqiad.wmnet' * 15:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1062.eqiad.wmnet' * 15:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 15:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 15:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 15:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 15:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1060.eqiad.wmnet' * 15:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 15:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 14:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1060.eqiad.wmnet' * 10:00 godog: switch cloudcephosd1049 to single nic - [[phab:T399180|T399180]] * 05:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1059.eqiad.wmnet' * 05:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1059.eqiad.wmnet' * 04:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1058.eqiad.wmnet' * 04:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1058.eqiad.wmnet' * 04:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1057.eqiad.wmnet' * 04:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1057.eqiad.wmnet' * 00:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' * 00:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1056.eqiad.wmnet' === 2025-11-17 === * 23:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1055.eqiad.wmnet' * 23:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1055.eqiad.wmnet' * 23:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1054.eqiad.wmnet' * 23:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' * 22:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1054.eqiad.wmnet' * 22:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' * 20:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' * 20:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' * 20:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1052.eqiad.wmnet' * 19:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1051.eqiad.wmnet' * 19:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1051.eqiad.wmnet' * 19:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1050.eqiad.wmnet' * 18:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1050.eqiad.wmnet' * 18:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1049.eqiad.wmnet' * 17:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1048.eqiad.wmnet' * 17:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' * 17:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047.eqiad.wmnet' * 17:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1046.eqiad.wmnet' * 17:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 17:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' * 16:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1045.eqiad.wmnet' * 15:58 godog: set ceph cluster back to out and rebalance - [[phab:T399180|T399180]] * 15:49 godog: set ceph cluster noout/norebalance and move cloudcephosd1048 to single nic - [[phab:T399180|T399180]] * 14:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1044.eqiad.wmnet' * 14:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 10:30 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch ([[phab:T409365|T409365]]) * 10:27 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch ([[phab:T409365|T409365]]) * 10:26 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch ([[phab:T409365|T409365]]) * 10:25 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch ([[phab:T409365|T409365]]) * 09:18 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,keystone * 09:18 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,keystone === 2025-11-16 === * 23:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1044.eqiad.wmnet' * 23:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 18:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1044.eqiad.wmnet' * 18:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 18:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1044.eqiad.wmnet' * 18:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 18:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' * 18:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1043.eqiad.wmnet' === 2025-11-15 === * 01:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' * 01:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1043.eqiad.wmnet' * 00:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1043.eqiad.wmnet' * 00:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1043.eqiad.wmnet' * 00:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1043.eqiad.wmnet' * 00:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1043.eqiad.wmnet' * 00:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 00:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 00:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 00:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 00:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 00:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance === 2025-11-14 === * 23:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' * 23:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1042.eqiad.wmnet' * 21:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1042.eqiad.wmnet' * 21:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1042.eqiad.wmnet' * 21:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1042.eqiad.wmnet' * 21:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1042.eqiad.wmnet' * 21:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1042.eqiad.wmnet' * 21:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1042.eqiad.wmnet' * 20:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' * 20:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1041.eqiad.wmnet' * 20:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' * 20:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' * 15:13 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/279 ([[phab:T409365|T409365]]) * 15:13 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/279 ([[phab:T409365|T409365]]) * 15:12 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/279 ([[phab:T409365|T409365]]) * 15:12 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/279 ([[phab:T409365|T409365]]) * 15:10 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch ([[phab:T409365|T409365]]) * 15:09 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch ([[phab:T409365|T409365]]) === 2025-11-13 === * 12:17 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:16 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:12 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.roll_reboot_cloudnets (exit_code=0) * 12:04 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.roll_reboot_cloudnets * 12:04 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,neutron * 11:52 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,neutron * 10:59 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/282 * 10:58 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/282 === 2025-11-11 === * 13:46 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:45 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-11-10 === * 15:46 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:45 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:44 taavi: rotate cookbook gitlab access token before(!) it expires [[phab:T409741|T409741]] * 15:37 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/280 * 15:36 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/280 * 15:35 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/280 * 15:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/280 * 15:33 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/280 * 15:33 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/280 * 14:53 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.roll_reboot_cloudnets (exit_code=0) * 14:43 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.roll_reboot_cloudnets * 14:28 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.roll_reboot_cloudnets (exit_code=99) * 14:28 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.roll_reboot_cloudnets * 14:27 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,neutron * 14:25 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,neutron === 2025-11-07 === * 11:33 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.vps.remove_instance (exit_code=1) for instance tools-db-7 * 11:33 fnegri@cloudcumin1001: START - Cookbook wmcs.vps.remove_instance for instance tools-db-7 === 2025-10-28 === * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 16:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 02:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 02:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-10-27 === * 20:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 20:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 18:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 18:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-10-25 === * 18:54 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) ([[phab:T405478|T405478]]) === 2025-10-24 === * 21:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 21:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 21:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 21:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 21:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 21:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 21:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 21:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 20:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2006-dev.codfw.wmnet' * 20:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2006-dev.codfw.wmnet' * 20:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 20:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 20:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 20:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2005-dev.codfw.wmnet' * 19:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2005-dev.codfw.wmnet' * 17:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2005-dev.codfw.wmnet' * 17:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2005-dev.codfw.wmnet' * 17:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2005-dev.codfw.wmnet' * 17:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2005-dev.codfw.wmnet' * 11:14 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T405478|T405478]]) * 08:21 filippo@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) ([[phab:T405478|T405478]]) === 2025-10-23 === * 21:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' * 21:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2004-dev.codfw.wmnet' * 21:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 21:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 21:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 21:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 21:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 21:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 21:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 21:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 21:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2004-dev.codfw.wmnet' * 20:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2004-dev.codfw.wmnet' * 16:02 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T405478|T405478]]) * 15:08 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T405478|T405478]]) * 07:08 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T405478|T405478]]) === 2025-10-22 === * 15:51 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T405478|T405478]]) * 07:51 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T405478|T405478]]) === 2025-10-21 === * 23:48 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T405478|T405478]]) * 15:47 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T405478|T405478]]) === 2025-10-20 === * 11:52 filippo@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 11:32 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-10-18 === * 05:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 05:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2025-10-16 === * 18:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 18:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 18:32 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 18:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 18:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 18:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-10-15 === * 14:22 filippo@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 14:13 filippo@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-10-14 === * 19:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 19:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 11:26 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::cloudweb<nowiki>}</nowiki>' * 11:20 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::cloudweb<nowiki>}</nowiki>' === 2025-10-08 === * 19:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 19:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-10-07 === * 18:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 18:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 18:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 18:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 18:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 18:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 18:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 18:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2025-10-01 === * 02:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 01:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 00:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 00:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 00:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' === 2025-09-30 === * 23:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1076.eqiad.wmnet<nowiki>}</nowiki>' * 23:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1076.eqiad.wmnet<nowiki>}</nowiki>' * 23:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' * 22:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' * 22:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1074.eqiad.wmnet<nowiki>}</nowiki>' * 22:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 22:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 22:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1074.eqiad.wmnet<nowiki>}</nowiki>' * 22:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1073.eqiad.wmnet<nowiki>}</nowiki>' * 22:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1073.eqiad.wmnet<nowiki>}</nowiki>' * 22:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1072.eqiad.wmnet<nowiki>}</nowiki>' * 22:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 22:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.eqiad.wmnet<nowiki>}</nowiki>' * 22:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.eqiad.wmnet<nowiki>}</nowiki>' * 22:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1072.eqiad.wmnet<nowiki>}</nowiki>' * 22:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1071.eqiad.wmnet<nowiki>}</nowiki>' * 21:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1071.eqiad.wmnet<nowiki>}</nowiki>' * 21:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1070.eqiad.wmnet<nowiki>}</nowiki>' * 21:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1070.eqiad.wmnet<nowiki>}</nowiki>' * 21:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1069.eqiad.wmnet<nowiki>}</nowiki>' * 21:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1069.eqiad.wmnet<nowiki>}</nowiki>' * 21:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1068.eqiad.wmnet<nowiki>}</nowiki>' * 21:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1068.eqiad.wmnet<nowiki>}</nowiki>' * 21:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' * 20:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' * 20:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' * 20:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' * 20:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' * 20:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' * 19:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' * 19:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' * 19:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' === 2025-09-26 === * 10:31 taavi: remove /var/lib/prometheus/node.d/kernel-messages.prom which got left as a leftover from the now-removed exporter === 2025-09-25 === * 23:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) === 2025-09-24 === * 15:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 09:27 volans@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' * 09:07 volans@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' * 08:53 volans@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' * 08:35 volans@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' * 08:35 volans@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 08:13 volans@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 07:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 02:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 01:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 01:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 01:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 01:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 01:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 01:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 01:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 01:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 00:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 00:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 00:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 00:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-09-23 === * 23:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 23:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 23:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 23:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 23:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 22:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' * 22:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' * 22:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' * 21:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' * 21:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' * 21:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' * 21:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 21:14 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' * 21:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' * 20:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 20:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' * 20:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 20:36 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 20:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 20:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' * 20:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' * 20:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' * 19:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' * 19:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 19:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' * 19:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' * 18:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' * 18:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 18:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 18:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:53 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.reactivate (exit_code=97) * 18:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 18:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 18:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' * 18:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' * 18:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' * 18:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 17:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 17:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' * 17:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet<nowiki>}</nowiki>' * 17:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet<nowiki>}</nowiki>' * 17:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1045.eqiad.wmnet<nowiki>}</nowiki>' * 17:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1045.eqiad.wmnet<nowiki>}</nowiki>' * 17:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' * 17:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' * 17:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 16:58 andrewbogott: test * 16:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 16:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 16:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 16:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:53 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 16:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 16:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 16:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 16:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 16:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 15:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 14:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:12 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 04:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 04:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 03:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 03:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 03:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 03:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 03:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 03:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 02:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 02:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 02:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 02:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 01:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-09-22 === * 22:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 22:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 22:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 22:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:25 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.reactivate (exit_code=97) * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 20:14 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 18:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 17:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 17:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 16:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 16:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 16:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 15:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-09-19 === * 00:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 00:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-09-18 === * 22:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 22:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 16:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:53 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 16:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 16:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-09-17 === * 22:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) === 2025-09-16 === * 14:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 14:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:32 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 13:40 andrewbogott: upgrading cloudcephosd1017 to bookworm/reef === 2025-09-15 === * 16:24 taavi: update nova-fullstack to run on trixie image === 2025-09-11 === * 20:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 19:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 19:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,designate * 19:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 19:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,designate * 19:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 19:22 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,designate * 19:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 19:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 19:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 19:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 19:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 15:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-09-10 === * 22:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 22:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 22:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 22:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 21:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 21:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 21:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:10 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 14:10 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 14:09 wmbot~dcaro@acme: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 14:09 wmbot~dcaro@acme: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 08:33 dhinus: add volans to cloud-vps domain admins: openstack role add --user volans --domain default --inherited admin === 2025-09-09 === * 16:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 13:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T402190|T402190]]) * 00:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2025-09-08 === * 14:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_mons (exit_code=0) * 13:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_mons === 2025-09-06 === * 11:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) === 2025-09-05 === * 21:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:46 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 15:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 15:46 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 15:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 15:43 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 15:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 15:41 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 15:41 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 15:41 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 15:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 15:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T401693|T401693]]) * 15:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 15:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T401693|T401693]]) * 15:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 15:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T401693|T401693]]) * 15:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 15:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T401693|T401693]]) * 15:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) === 2025-09-04 === * 16:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T395910|T395910]]) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T395910|T395910]]) === 2025-08-29 === * 19:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 14:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 05:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 00:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2025-08-28 === * 20:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 12:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 12:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 09:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 07:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 07:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 04:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 04:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 01:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) === 2025-08-27 === * 22:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 22:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 12:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 06:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 03:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 03:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 01:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:38 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=97) * 01:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 01:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) === 2025-08-26 === * 20:57 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:35 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T401693|T401693]]) * 01:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 01:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T401693|T401693]]) * 01:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) === 2025-08-25 === * 21:17 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) ([[phab:T401693|T401693]]) * 21:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 17:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 17:40 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 17:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 17:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 17:37 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 17:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 17:35 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 17:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 17:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 14:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 13:36 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) ([[phab:T401693|T401693]]) * 04:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 04:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T401693|T401693]]) === 2025-08-24 === * 17:39 andrewbogott: removing /var/lib/prometheus/node.d/check_disk_space.prom on all cloudvirts and a few other servers. I am taking this file to be obsolete after https://gerrit.wikimedia.org/r/c/operations/puppet/+/1180501/2 ; removing it seems to clear the alert. * 17:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 10:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T401693|T401693]]) === 2025-08-23 === * 19:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 19:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T401693|T401693]]) * 04:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 04:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T401693|T401693]]) * 04:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) === 2025-08-22 === * 20:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 14:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 14:29 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 14:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 14:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T395910|T395910]]) * 14:24 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T395910|T395910]]) * 14:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T395910|T395910]]) * 14:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T395910|T395910]]) * 08:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) === 2025-08-21 === * 16:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 15:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:57 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 15:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 15:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 15:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 14:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 14:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 14:41 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 14:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 14:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:38 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 14:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:36 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 14:35 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 12:58 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T402499|T402499]]) * 12:57 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T402499|T402499]]) * 12:50 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T402499|T402499]]) * 12:18 dcaro: destroying osd 66 on cloudcephosd1004, will recreate ([[phab:T402499|T402499]]) * 12:17 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T402499|T402499]]) * 10:35 dcaro: starting ceph-osd@69 on cloudcephosd1004 ([[phab:T402499|T402499]]) * 10:32 dcaro: starting ceph-osd@68 on cloudcephosd1004 ([[phab:T402499|T402499]]) * 09:44 dcaro: starting ceph-osd@66 on cloudcephosd1004 ([[phab:T402499|T402499]]) * 09:22 dcaro: test fol sal === 2025-08-20 === * 13:21 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 13:21 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 13:20 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 13:20 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 13:20 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 13:20 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 13:19 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 13:19 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 13:19 filippo@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 13:19 filippo@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 09:05 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.reboot_node (exit_code=0) ([[phab:T401319|T401319]]) * 09:01 fnegri@cloudcumin1001: START - Cookbook wmcs.ceph.reboot_node ([[phab:T401319|T401319]]) === 2025-08-18 === * 16:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_mons (exit_code=0) * 15:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_mons ([[phab:T402190|T402190]]) * 15:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 15:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T402190|T402190]]) * 15:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 15:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T402190|T402190]]) === 2025-08-15 === * 20:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T401693|T401693]]) * 20:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 19:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T401693|T401693]]) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) * 19:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T401693|T401693]]) * 19:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T401693|T401693]]) === 2025-08-13 === * 23:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 23:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 23:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 23:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 23:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 22:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 22:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 21:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 19:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 19:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 19:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 18:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 18:43 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds === 2025-08-12 === * 19:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 19:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-08-07 === * 09:52 taavi: remove ssh key from uid=soni LDAP user [[phab:T401318|T401318]] === 2025-08-06 === * 13:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-07-29 === * 08:08 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 08:05 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-07-24 === * 20:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 20:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 20:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,nova * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,nova * 15:44 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:42 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:57 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/258 * 14:57 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/258 * 14:57 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/258 * 14:56 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/258 * 13:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,nova * 13:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,nova * 11:01 dcaro: stopping all ceph osds in codfw1 to avoid spamming the mons ([[phab:T400334|T400334]]) * 07:37 dcaro: downgrading the codfw1 ceph mons to pacific, to do a rebuild instead of in-place upgrade to quincy === 2025-07-22 === * 19:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_mons (exit_code=0) * 19:40 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_mons * 19:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 19:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 18:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 18:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 14:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 14:40 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:36 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.upgrade_osds (exit_code=97) * 14:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 10:06 wmbot~dcaro@hephaestus: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 10:04 wmbot~dcaro@hephaestus: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 10:04 wmbot~dcaro@hephaestus: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 10:03 wmbot~dcaro@hephaestus: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 10:03 wmbot~dcaro@hephaestus: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 07:45 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T399870|T399870]]) * 07:44 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T399870|T399870]]) * 00:38 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) === 2025-07-21 === * 20:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 20:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:20 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 20:20 dcaro@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 20:08 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:07 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 20:07 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:07 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 20:07 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 20:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:44 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) ([[phab:T399858|T399858]]) * 17:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 17:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 17:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 17:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 17:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 17:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 16:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 16:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 16:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1006.eqiad.wmnet' ([[phab:T395255|T395255]]) * 16:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1006.eqiad.wmnet' ([[phab:T395255|T395255]]) * 16:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1005.eqiad.wmnet' ([[phab:T395255|T395255]]) * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1005.eqiad.wmnet' ([[phab:T395255|T395255]]) * 15:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T399858|T399858]]) * 15:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T399858|T399858]]) * 15:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T399858|T399858]]) * 15:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T399858|T399858]]) * 15:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T399858|T399858]]) * 15:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T399858|T399858]]) * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T399858|T399858]]) * 15:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T399858|T399858]]) * 13:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T399858|T399858]]) === 2025-07-20 === * 10:44 andrewbogott: rebooting cloudcephosd1006 to give us another few days before the memory runs out. [[phab:T399858|T399858]] === 2025-07-19 === * 13:05 andrewbogott: restarted neutron-metadata-agent on cloudnet100[56], again === 2025-07-17 === * 16:02 dcaro: restart cloudcephosd1006 osd 45 for the memory limit take effect * 15:57 dcaro: lowered memory target for cloudcephosd1006 to 5G === 2025-07-16 === * 18:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 18:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:32 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 09:47 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1073.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T399212|T399212]]) * 09:44 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1073.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T399212|T399212]]) === 2025-07-15 === * 23:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 23:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-07-11 === * 18:38 andrewbogott: it didn't * 18:32 andrewbogott: rebooting cloudceph1013 to see if its missing OSD drive reappears * 18:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 18:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 18:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 18:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 17:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 17:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 16:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for service: project,designate * 16:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 15:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 15:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 15:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 15:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 15:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 00:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 00:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-07-10 === * 23:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 22:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 22:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 22:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:03 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.reactivate (exit_code=97) * 15:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 12:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 12:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 12:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 12:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 04:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 04:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-07-09 === * 23:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 18:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds ([[phab:T306820|T306820]]) * 15:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_mons (exit_code=0) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_mons ([[phab:T306820|T306820]]) * 14:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 14:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 14:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 13:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 13:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 01:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1073.eqiad.wmnet' ([[phab:T394333|T394333]]) * 01:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1073.eqiad.wmnet' ([[phab:T394333|T394333]]) === 2025-07-08 === * 23:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 23:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 21:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 21:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 20:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 20:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 20:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 20:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 20:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 20:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 20:35 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 20:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 20:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 20:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 19:57 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:24 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:12 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 19:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 19:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:45 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:45 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 18:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 18:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 17:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 17:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 16:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.reactivate (exit_code=0) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 15:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 15:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.reactivate (exit_code=99) * 14:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate * 14:58 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.reactivate (exit_code=97) * 14:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.reactivate === 2025-07-07 === * 21:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 21:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 21:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 21:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 21:06 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=97) * 21:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy === 2025-07-03 === * 14:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,designate * 14:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate * 14:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' * 14:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' * 13:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2006-dev.codfw.wmnet' * 13:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 13:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2005-dev.codfw.wmnet' * 13:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 13:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2006-dev.codfw.wmnet' * 13:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' === 2025-07-02 === * 19:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.upgrade_osds (exit_code=99) * 18:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 18:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_ceph_node (exit_code=0) * 18:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_ceph_node * 18:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_ceph_node (exit_code=0) * 18:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_ceph_node * 17:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 17:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_ceph_node (exit_code=0) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_ceph_node * 16:29 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.upgrade_ceph_node (exit_code=97) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_ceph_node * 16:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:52 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/253 * 12:52 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/253 === 2025-07-01 === * 20:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.reboot_node (exit_code=99) * 20:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.reboot_node * 18:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 18:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 16:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 15:57 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:17 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:16 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-06-30 === * 19:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 * 19:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 === 2025-06-26 === * 17:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 17:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 17:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 17:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 15:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1003.eqiad.wmnet<nowiki>}</nowiki>' * 15:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1003.eqiad.wmnet<nowiki>}</nowiki>' * 15:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1002.eqiad.wmnet<nowiki>}</nowiki>' * 15:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1002.eqiad.wmnet<nowiki>}</nowiki>' * 15:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1001.eqiad.wmnet<nowiki>}</nowiki>' * 14:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1001.eqiad.wmnet<nowiki>}</nowiki>' * 14:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 08:53 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 08:52 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:51 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/251 * 08:51 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/251 === 2025-06-25 === * 21:22 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 21:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 21:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 21:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 21:10 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 21:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:35 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 20:35 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:35 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 20:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 20:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 20:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 20:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 20:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:29 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 20:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:28 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 20:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:27 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 20:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-06-24 === * 21:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 21:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 20:50 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 20:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 20:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 20:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2025-06-23 === * 22:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,neutron * 21:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,neutron * 21:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,neutron * 21:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,neutron * 19:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,neutron * 19:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,neutron * 19:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,neutron * 19:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,neutron * 13:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,neutron * 13:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,neutron * 13:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,neutron * 13:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,neutron === 2025-06-21 === * 16:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 16:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 03:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 03:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2025-06-20 === * 20:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova,cinder,neutron * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova,cinder,neutron * 17:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' === 2025-06-19 === * 04:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 04:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2025-06-18 === * 20:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 20:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 17:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' === 2025-06-17 === * 20:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 20:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,cinder * 20:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,cinder * 20:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 19:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 19:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds * 19:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.upgrade_osds (exit_code=0) * 18:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.upgrade_osds === 2025-06-14 === * 22:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 22:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 20:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 19:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 18:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 18:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 18:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T309789|T309789]]) * 16:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 13:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 13:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1023.eqiad.wmnet' ([[phab:T394727|T394727]]) * 13:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1023.eqiad.wmnet' ([[phab:T394727|T394727]]) * 13:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 13:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 13:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 13:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 13:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 13:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 13:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 12:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 12:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 12:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 12:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 11:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,heat * 11:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,heat * 07:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T309789|T309789]]) * 05:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 05:14 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 05:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 05:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 05:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 05:12 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 04:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 04:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T309789|T309789]]) * 04:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 04:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 04:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 04:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 04:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 03:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 03:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 03:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T309789|T309789]]) * 03:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 03:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T309789|T309789]]) * 03:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 03:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 03:54 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 03:51 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) ([[phab:T309789|T309789]]) * 02:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 * 02:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 * 02:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 02:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 01:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 01:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 01:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 01:36 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 00:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T309789|T309789]]) === 2025-06-13 === * 21:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 21:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 21:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 21:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:22 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 20:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 19:12 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:12 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 19:12 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:12 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=97) * 19:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 19:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 19:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 19:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 19:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 19:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 18:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T309789|T309789]]) * 14:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:03 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 13:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 11:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 09:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 06:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T309789|T309789]]) * 04:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 04:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:01 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 03:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 02:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 02:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 02:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 02:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 02:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 02:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 02:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 02:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 00:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 00:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 00:57 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) ([[phab:T309789|T309789]]) * 00:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 00:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) === 2025-06-12 === * 23:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 * 23:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 * 23:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 * 23:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/248 * 21:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 21:40 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 21:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 21:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:40 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 20:35 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 20:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 20:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 20:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T396363|T396363]]) * 19:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T396363|T396363]]) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 19:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T309789|T309789]]) * 19:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T396363|T396363]]) * 19:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T396363|T396363]]) * 19:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 17:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T309789|T309789]]) * 17:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T396363|T396363]]) * 16:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 15:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 14:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:58 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 14:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 14:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:25 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:24 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T396363|T396363]]) * 11:41 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) * 05:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T309789|T309789]]) * 02:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T309789|T309789]]) === 2025-06-11 === * 21:58 andrewbogott: created new keystone role in eqiad1 and codfw1dev, 'object_storage' [[phab:T396594|T396594]] * 16:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,cinder * 16:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,cinder * 15:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 15:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 15:20 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 14:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,cinder * 14:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,cinder === 2025-06-10 === * 16:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 16:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 16:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 16:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/246 * 15:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/246 * 15:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/246 * 15:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/246 * 15:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/246 * 15:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/246 * 15:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 15:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/245 * 10:16 taavi: add cloud-private addresses to eqiad hosts [[phab:T379283|T379283]] === 2025-06-09 === * 20:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 20:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 19:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 19:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 19:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:13 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment codfw1dev for service: project,designate * 19:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 19:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,designate * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 19:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,designate * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 19:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,keystone * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,keystone * 19:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 19:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services * 19:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 13:14 taavi: add AAAA record to openstack.codfw1dev.wikimediacloud.org [[phab:T379282|T379282]] * 13:13 taavi: add AAAA record to openstack.codfw1dev.wikimediacloud.org [[phab:T347148|T347148]] === 2025-06-07 === * 19:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 18:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 18:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova * 18:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova === 2025-06-06 === * 20:59 andrewbogott: restarting all designate services on all cloudcontrols in eqiad1 === 2025-06-05 === * 20:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 17:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 17:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 17:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,octavia * 17:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,octavia * 17:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,octavia * 17:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,octavia * 15:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 14:52 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 14:41 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2025-06-03 === * 22:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 22:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 22:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/244 * 22:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/244 * 22:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/244 * 22:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/244 * 20:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 20:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 19:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 19:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 18:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:44 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 18:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:43 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 18:43 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T395910|T395910]]) * 18:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T395910|T395910]]) * 17:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T395910|T395910]]) * 17:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T395910|T395910]]) * 17:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T395910|T395910]]) * 17:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T395910|T395910]]) * 17:36 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 17:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-06-02 === * 21:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 21:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 21:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/237 * 21:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/237 * 17:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 17:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/242 * 17:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/242 * 16:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/241 * 16:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/241 * 16:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/241 * 16:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/241 * 16:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/241 * 16:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/241 * 16:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 16:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/240 * 15:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/239 * 15:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/239 * 15:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/239 * 15:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/239 * 00:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T394333|T394333]]) === 2025-06-01 === * 18:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T394333|T394333]]) === 2025-05-31 === * 23:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services ([[phab:T395742|T395742]]) * 23:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services ([[phab:T395742|T395742]]) * 23:38 andrewbogott: failing over from cloudnet1005 to 1006 in hopes of unsticking [[phab:T395742|T395742]] * 23:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,neutron ([[phab:T395742|T395742]]) * 23:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,neutron ([[phab:T395742|T395742]]) === 2025-05-30 === * 14:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-05-28 === * 22:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1076.eqiad.wmnet' ([[phab:T390914|T390914]]) * 22:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1076.eqiad.wmnet' ([[phab:T390914|T390914]]) * 22:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1075.eqiad.wmnet' ([[phab:T390914|T390914]]) * 22:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1075.eqiad.wmnet' ([[phab:T390914|T390914]]) * 22:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1074.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1074.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1073.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1073.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1072.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1072.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1071.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1070.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1070.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1069.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1069.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1068.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1068.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T390914|T390914]]) * 21:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 20:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 19:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1002-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1002-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1001-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1001-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1004.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1004.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T390914|T390914]]) * 19:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:03 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 17:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 17:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1006.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1006.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1005.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1005.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudrabbit1001.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudrabbit1001.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudrabbit1002.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudrabbit1002.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudrabbit1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudrabbit1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudrabbot1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudrabbot1003.eqiad.wmnet' ([[phab:T390914|T390914]]) * 17:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T390914|T390914]]) * 16:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T390914|T390914]]) * 16:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T390914|T390914]]) * 16:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T390914|T390914]]) * 16:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T390914|T390914]]) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:32 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 15:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 15:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 10:40 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,designate * 10:39 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate === 2025-05-26 === * 13:13 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=0) for host cloudnet2006-dev.codfw.wmnet * 13:10 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.reboot_node for host cloudnet2006-dev.codfw.wmnet * 13:06 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=0) for host cloudnet2005-dev.codfw.wmnet * 13:03 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.reboot_node for host cloudnet2005-dev.codfw.wmnet * 12:54 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.roll_reboot_cloudnets (exit_code=99) * 12:54 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.roll_reboot_cloudnets * 12:49 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.roll_reboot_cloudnets (exit_code=99) * 12:49 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.roll_reboot_cloudnets * 12:48 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::control<nowiki>}</nowiki>' * 12:36 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::control<nowiki>}</nowiki>' * 12:31 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 11:49 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 11:33 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 11:27 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 11:22 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 11:21 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 11:18 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 11:17 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::codfw1dev::nova::compute::service<nowiki>}</nowiki>' * 07:52 taavi: add new /dev/sdb back to software raid after it was replaced in cloudcephmon1004 === 2025-05-23 === * 09:00 dhinus: failover dumps_dist_active_vps to clouddumps1001 ([[phab:T383723|T383723]]) === 2025-05-20 === * 19:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 15:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 14:57 andrewbogott: resetting eqiad1 rabbitmq in an attempt to resolve [[phab:T394790|T394790]] * 03:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,cinder * 03:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,cinder * 02:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,cinder * 02:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,cinder * 00:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 00:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) === 2025-05-19 === * 22:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 22:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 22:43 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for service: project,nova * 22:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova * 22:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova * 22:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova * 22:26 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for service: project,nova * 22:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova * 22:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 21:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 21:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 21:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 21:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 21:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 21:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 20:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 19:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 18:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 18:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 18:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 18:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 18:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 18:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 18:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 18:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T394727|T394727]]) * 18:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T394727|T394727]]) * 17:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' * 17:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1075.eqiad.wmnet<nowiki>}</nowiki>' * 17:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1072.eqiad.wmnet<nowiki>}</nowiki>' * 17:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1072.eqiad.wmnet<nowiki>}</nowiki>' * 17:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1076.eqiad.wmnet<nowiki>}</nowiki>' * 17:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1076.eqiad.wmnet<nowiki>}</nowiki>' * 17:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1073.eqiad.wmnet<nowiki>}</nowiki>' * 17:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1073.eqiad.wmnet<nowiki>}</nowiki>' * 15:24 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:23 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-05-15 === * 15:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/235 * 15:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/235 * 15:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/235 * 15:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/235 * 14:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/234 * 14:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/234 * 14:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/234 * 14:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/234 * 13:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-05-14 === * 17:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 17:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 17:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/231 * 17:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/231 * 12:59 taavi: powercycle unresponsive cloudnet2006-dev === 2025-05-12 === * 20:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 19:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 08:17 taavi: powercycle clouservices2005-dev.codfw.wmnet * 03:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 03:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 02:36 andrewbogott: rebooting cloudnet2005-dev from mgmt -- ssh is failing and the console shows a user prompt but not a password prompt. === 2025-05-11 === * 13:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,nova * 13:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,nova === 2025-05-07 === * 20:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1002-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1002-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1001-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1001-dev.eqiad.wmnet' ([[phab:T390914|T390914]]) * 20:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2006-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 20:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2006-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 20:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 20:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 20:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 20:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,designate * 19:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 19:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2004-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2006-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2006-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 19:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2004-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2004-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004-dev.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2004.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004.codfw.wmnet' ([[phab:T390914|T390914]]) * 18:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2004.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004.eqiad.wmnet' ([[phab:T390914|T390914]]) * 18:19 andrewbogott: upgrading codfw1dev to version 'epoxy' [[phab:T390914|T390914]] * 18:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=0) ([[phab:T390914|T390914]]) * 18:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T390914|T390914]]) * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,heat * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,heat * 16:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,magnum * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,magnum * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,magnum,heat * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,magnum,heat * 16:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,magnum,heat * 16:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,magnum,heat * 16:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,heat * 16:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,heat * 16:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,heat * 16:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,heat * 16:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,magnum * 16:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,magnum * 15:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,nova * 15:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,nova * 15:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,nova * 15:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,nova * 15:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,nova * 15:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,nova * 15:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,nova * 15:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,nova * 13:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,neutron * 13:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,neutron * 13:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,neutron * 13:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,neutron * 12:48 taavi: updating all security group rules referencing old 172.16.0.0/21 subnet to reference new ip space instead ([[phab:T379175|T379175]]) * 10:49 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:48 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 00:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 00:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-05-06 === * 03:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) === 2025-05-05 === * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 12:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 10:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 08:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 03:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 00:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 00:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 00:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 00:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 00:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 00:06 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 00:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 00:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 00:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 00:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 00:03 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 00:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 00:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 00:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-05-04 === * 23:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 23:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 22:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 22:43 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 22:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 22:41 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 22:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 22:40 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 22:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 22:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 22:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 22:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 21:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 21:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 21:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 21:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 21:31 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 21:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 21:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 21:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2025-05-03 === * 02:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T393196|T393196]]) === 2025-05-02 === * 23:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T393196|T393196]]) * 21:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T393196|T393196]]) * 18:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T393196|T393196]]) * 18:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T393196|T393196]]) * 16:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T393196|T393196]]) * 16:20 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=97) ([[phab:T393196|T393196]]) * 16:14 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T393196|T393196]]) * 16:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 16:14 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:09 andrewbogott: sudo cumin --force O<nowiki>{</nowiki>*<nowiki>}</nowiki> "dpkg --list {{!}} grep puppetserver && systemctl restart puppetserver.service" # work around a package update that is causing some puppetservers to error out === 2025-04-29 === * 13:01 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:01 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:49 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/225 * 12:48 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/225 * 11:48 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:47 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:15 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:14 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:13 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:13 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:11 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:11 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:10 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:10 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:09 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:09 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:05 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 11:05 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 10:55 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 10:55 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 10:52 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 10:51 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/224 * 08:16 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 08:15 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 07:46 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/223 * 07:45 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/223 * 07:43 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 07:43 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 07:42 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 07:41 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 07:41 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/221 * 07:40 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/221 === 2025-04-28 === * 15:49 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/221 * 15:49 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/221 * 11:25 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:21 taavi: migrating opentofu managed default security group rules [[phab:T392799|T392799]] * 11:20 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:13 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 08:13 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-04-26 === * 07:48 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) ([[phab:T390134|T390134]]) * 07:48 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.unset_cluster_maintenance ([[phab:T390134|T390134]]) * 07:48 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T390134|T390134]]) * 07:47 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T390134|T390134]]) === 2025-04-25 === * 12:55 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/218 * 12:54 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/218 * 12:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:30 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 04:29 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T390134|T390134]]) === 2025-04-24 === * 17:25 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T390134|T390134]]) * 13:29 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:29 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 07:56 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for service: project,designate * 07:55 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate === 2025-04-23 === * 21:45 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 21:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 21:45 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 21:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 21:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 21:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 21:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 21:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 21:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 21:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 21:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 21:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 21:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 21:21 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 21:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 19:15 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 19:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:03 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:32 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 16:31 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 16:29 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 16:29 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 16:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:17 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 16:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:29 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:28 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:27 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:27 arturo: enabling IPv6 dualstack on neutron virtual router ([[phab:T380174|T380174]]) * 14:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:22 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/217 * 14:22 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/217 * 12:15 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:15 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 12:12 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:12 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 12:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:04 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 12:03 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:54 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:54 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:53 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:49 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:45 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:45 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:44 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:42 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:14 arturo: enable IPv6 on cloudgw ([[phab:T380174|T380174]]) -- includes server reboot === 2025-04-22 === * 04:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) === 2025-04-21 === * 23:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 23:34 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 23:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 23:18 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 23:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 22:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 22:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 22:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 22:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 22:09 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 22:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 21:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 16:43 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 16:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 11:05 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 01:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 01:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-04-16 === * 16:49 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:49 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,cinder * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,cinder * 13:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,cinder * 13:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,cinder * 12:34 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:33 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:26 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:26 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:26 arturo: merging network change in neutron https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/198 * 10:10 dcaro: upgrade spicerack on cloudcumin2001 to 10.1.0 * 10:08 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 10:07 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 10:07 dcaro: upgrade spicerack on cloudcumin1001 to 10.1.0 === 2025-04-15 === * 15:17 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:16 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:16 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:15 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:13 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:49 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:49 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:38 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:37 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:31 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:30 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-04-14 === * 17:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 17:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 16:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:50 andrewbogott: granting 'tofuadmin' user inherited 'member' role in all projects, this should fix some policy mishaps * 16:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-04-13 === * 22:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 22:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-04-11 === * 11:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:50 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:50 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:50 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:49 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:07 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:06 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-04-10 === * 12:38 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:05 arturo: [codfw1dev] root@cloudcontrol2004-dev:~# wmcs-makedomain --project testlabs --domain testlabs.codfw1dev.wmcloud.org --orig-project cloudinfra-codfw1dev ([[phab:T391325|T391325]]) === 2025-04-09 === * 23:23 bd808: Rebooting tools-sgebastion-10 (login-buster.toolforge.org) for high load/unresponsive NFS mounts ([[phab:T391538|T391538]]) * 10:18 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:17 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:06 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:39 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 08:38 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 04:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 03:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 03:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for all services * 03:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 03:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for all services * 03:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 03:30 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 03:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 03:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for all services * 03:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 03:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for all services * 03:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2025-04-08 === * 22:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1011.eqiad.wmnet' * 22:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' * 22:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol1011.eqiad.wmnet' * 22:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1011.eqiad.wmnet' * 11:14 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2025-04-07 === * 23:43 bd808: `sudo service maintain-dbusers restart` on cloudcontrol1007 after reports of missing replica.my.cnf and finding the journal for the service empty. * 15:30 arturo: [codfw1dev] testlabs create a bunch of VMs by hand, like `networktests-vlan-legacy-floating` [[phab:T380728|T380728]] * 15:21 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:20 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 14:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 14:06 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:58 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:30 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:29 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:24 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:24 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:23 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:23 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-04-06 === * 02:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 02:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-04-04 === * 15:14 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:13 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 05:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 05:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 05:10 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment codfw1dev for all services * 05:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 05:04 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment codfw1dev for all services * 05:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services === 2025-04-03 === * 22:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 22:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 22:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T381499|T381499]]) * 22:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T381499|T381499]]) * 21:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T381499|T381499]]) * 21:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 21:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 21:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T381499|T381499]]) * 21:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T381499|T381499]]) * 21:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 20:12 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 20:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 20:11 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) on deployment eqiad1 for all services * 20:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 19:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment eqiad1 for all services * 19:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 19:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudrabbit1003.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudrabbit1003.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudrabbit1002.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudrabbit1002.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudrabbit1001.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudrabbit1001.eqiad.wmnet' ([[phab:T381499|T381499]]) * 13:29 taavi: run wmcs-wikireplica-dns to create transitional x3 CNAMEs [[phab:T390954|T390954]] === 2025-04-02 === * 20:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T381499|T381499]]) * 20:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T381499|T381499]]) * 19:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup2004.codfw.wmnet' ([[phab:T381499|T381499]]) * 18:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1004.eqiad.wmnet' ([[phab:T381499|T381499]]) * 18:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1003.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1003.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T381499|T381499]]) * 17:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.unset_maintenance (exit_code=0) ({{Gerrit|1133432}}) * 16:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.unset_maintenance ({{Gerrit|1133432}}) * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:22 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.tofu (exit_code=97) running tofu plan+apply for main branch * 16:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:05 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 16:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 15:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1007.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1007.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:32 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:26 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:20 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:18 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T381499|T381499]]) * 14:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 13:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T381499|T381499]]) * 13:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=0) ([[phab:T381499|T381499]]) * 13:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T381499|T381499]]) * 09:53 dhinus: systemctl restart maintain-dbusers.service (attempting to fix a weird issue) === 2025-04-01 === * 21:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for service: project,designate * 21:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 21:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 21:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 21:00 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.tofu (exit_code=97) running tofu plan+apply for main branch * 21:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:50 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:50 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-03-31 === * 15:54 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) * 15:54 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.unset_cluster_maintenance * 15:54 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) * 15:54 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.unset_cluster_maintenance * 15:53 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=0) * 15:53 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.set_cluster_in_maintenance === 2025-03-27 === * 14:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 14:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 12:00 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T390134|T390134]]) * 08:41 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T390134|T390134]]) * 08:41 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T390134|T390134]]) * 08:41 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T390134|T390134]]) === 2025-03-26 === * 10:38 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-03-25 === * 14:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:22 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:21 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:14 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-03-24 === * 11:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-03-22 === * 03:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for service: project,designate * 03:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for service: project,designate === 2025-03-19 === * 14:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,designate * 14:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,designate * 12:02 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for service: project,task_id,no_dologmsg,cluster_name,all_services,nova,glance,keystone,cinder,neutron,trove,magnum,heat,swift,designate,filter_nodes * 12:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for service: project,task_id,no_dologmsg,cluster_name,all_services,nova,glance,keystone,cinder,neutron,trove,magnum,heat,swift,designate,filter_nodes * 06:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 06:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 06:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 06:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1002-dev.eqiad.wmnet' * 06:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 06:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2006-dev.codfw.wmnet' * 06:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1002-dev.eqiad.wmnet' * 06:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1001-dev.eqiad.wmnet' * 06:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2006-dev.codfw.wmnet' * 06:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2005-dev.codfw.wmnet' * 06:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1001-dev.eqiad.wmnet' * 05:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2006-dev.codfw.wmnet' * 05:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2005-dev.codfw.wmnet' * 05:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' * 05:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 05:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2006-dev.codfw.wmnet' * 05:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 05:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2006-dev.codfw.wmnet' * 05:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 05:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2006-dev.codfw.wmnet' * 05:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 05:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2004-dev.codfw.wmnet' * 05:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2006-dev.codfw.wmnet' * 05:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2005-dev.codfw.wmnet' * 05:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 05:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2005-dev.codfw.wmnet' * 05:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:32 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services * 05:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services * 05:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:22 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudnet2005-dev.codfw.wmnet' * 05:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:12 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) on host 'cloudnet2005-dev.codfw.wmnet' * 05:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 05:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1001-dev.eqiad.wmnet' * 05:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 05:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 05:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1001-dev.eqiad.wmnet' * 04:58 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 04:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 04:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2006-dev.codfw.wmnet' * 04:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 04:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 04:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2004-dev.codfw.wmnet' * 04:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004-dev.codfw.wmnet' * 04:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' * 04:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2004-dev.codfw.wmnet' * 04:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 04:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' * 04:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' * 04:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 04:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 04:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 04:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 04:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 04:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 03:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 02:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 02:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 02:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 02:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=0) ([[phab:T381499|T381499]]) * 02:12 andrewbogott: upgrading codfw1dev to openstack 'dalmation' [[phab:T381499|T381499]] * 02:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T381499|T381499]]) === 2025-03-14 === * 13:36 volans: installed cumin v5.1.1 on cloudcumin* hosts === 2025-03-05 === * 17:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 17:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services * 00:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment eqiad1 for all services * 00:43 andrewbogott: restarting all openstack services in hopes of gettting dns unstuck * 00:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack on deployment eqiad1 for all services === 2025-03-04 === * 08:52 arturo: [[phab:T387828|T387828]] depooled galera on cloudcontrol1005 === 2025-03-01 === * 19:35 andrewbogott: installed new bookworm base images in eqiad1 * 19:35 andrewbogott: installed new bookworm and bullseye base images in codfw1dev === 2025-02-28 === * 15:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2025-02-27 === * 15:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2025-02-24 === * 00:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 00:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 00:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 00:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 00:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 00:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 00:06 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.tofu (exit_code=97) running tofu plan for main branch * 00:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2025-02-23 === * 21:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2025-02-21 === * 12:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2025-02-20 === * 17:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T386083|T386083]]) * 17:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T386083|T386083]]) === 2025-02-19 === * 13:26 arturo: manual failover of cloudgw1004 to cloudgw1003 [[phab:T382356|T382356]] * 01:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 01:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 00:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2025-02-18 === * 12:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-02-13 === * 13:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2025-02-11 === * 11:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T386083|T386083]]) * 11:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T386083|T386083]]) * 11:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1047' ([[phab:T386083|T386083]]) * 11:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047' ([[phab:T386083|T386083]]) === 2025-02-07 === * 13:52 root@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=0) for host cloudnet1005.eqiad.wmnet ([[phab:T384946|T384946]]) * 13:48 root@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.reboot_node for host cloudnet1005.eqiad.wmnet ([[phab:T384946|T384946]]) * 02:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 02:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 01:48 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 01:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 01:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 01:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 01:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' * 01:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' * 01:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1037.eqiad.wmnet<nowiki>}</nowiki>' * 01:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1037.eqiad.wmnet<nowiki>}</nowiki>' * 01:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1038.eqiad.wmnet<nowiki>}</nowiki>' * 00:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1038.eqiad.wmnet<nowiki>}</nowiki>' * 00:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1039.eqiad.wmnet<nowiki>}</nowiki>' * 00:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1039.eqiad.wmnet<nowiki>}</nowiki>' === 2025-02-06 === * 21:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 21:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 21:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 21:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' * 21:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 21:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 20:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' * 20:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' * 20:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 20:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 20:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 20:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' * 20:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' * 20:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' * 20:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' * 20:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' * 20:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' * 20:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1045.eqiad.wmnet<nowiki>}</nowiki>' * 19:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' * 19:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 19:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1045.eqiad.wmnet<nowiki>}</nowiki>' * 19:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet<nowiki>}</nowiki>' * 19:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 19:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' * 19:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet<nowiki>}</nowiki>' * 19:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' * 19:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' * 19:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' * 19:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' * 19:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' * 19:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' * 19:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' * 18:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' * 18:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' * 18:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' * 18:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 18:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 18:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 18:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' * 18:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' * 17:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 17:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' * 17:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' * 17:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' * 17:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' * 17:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' * 17:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' * 17:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' * 17:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' * 16:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' * 16:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' * 16:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' * 16:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' * 16:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' * 13:45 andrewbogott: cold-migrating all remaining VMs in [[phab:T385264|T385264]] except for 'integration' and 'tools' VMs === 2025-02-05 === * 19:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 19:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance === 2025-02-04 === * 12:48 arturo: replacing cloudgw1002 with cloudgw1004 - [[phab:T382356|T382356]] * 09:22 arturo: fleet-wide restart of puppetservers [[phab:T385553|T385553]] === 2025-02-03 === * 13:42 andrewbogott: rebooting proxy-04.project-proxy for the ceph OSD mishap === 2025-02-02 === * 17:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 17:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 17:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 17:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 17:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 17:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' * 17:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2004-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 16:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 16:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 16:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 16:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 16:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance === 2025-01-31 === * 01:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2025-01-30 === * 21:39 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 21:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 20:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:46 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 14:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 14:33 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 14:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 13:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 13:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 13:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 13:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 13:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 13:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 13:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 13:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 13:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 12:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1035.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1035.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1034.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1034.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1033.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1033.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1032.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1032.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1031.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1031.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 04:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 03:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 02:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 02:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 00:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 00:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) === 2025-01-29 === * 23:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 23:43 andrewbogott: resetting rabbitmq in eqiad1 in hopes that it will resolve mysterious openstack misbehaviors * 20:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2005-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 20:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 20:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 19:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 19:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 19:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 19:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt2006-dev.codfw.wmnet<nowiki>}</nowiki>' * 18:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 18:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 18:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2009-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 18:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2009-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2006-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2006-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2004-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1007.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2004-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2005-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1007.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1006.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol2005-dev.codfw.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1006.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1005.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 17:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1005.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T384946|T384946]]) * 12:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-01-28 === * 21:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 21:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 21:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 21:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 20:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 16:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 13:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 13:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 13:02 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) ([[phab:T348643|T348643]]) * 13:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 02:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 00:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 00:49 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) ([[phab:T348643|T348643]]) * 00:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) === 2025-01-27 === * 21:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) * 21:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.unset_cluster_maintenance * 21:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 21:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 21:52 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 21:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 14:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 14:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 14:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 14:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 14:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 14:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 14:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 14:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-01-24 === * 13:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 13:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node === 2025-01-23 === * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 20:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 19:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 18:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 18:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 17:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 15:49 dhinus: cumin 'P:base::cloud_production' 'rm /var/lib/prometheus/node.d/kernel-panic.prom' [[phab:T382961|T382961]] * 15:32 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 15:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 14:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 14:57 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 13:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2025-01-22 === * 14:52 dhinus: cloudcumin[12]001 upgrade spicerack from 8.15.2 to 9.1.0 === 2025-01-21 === * 13:42 andrewbogott: migrating/rebooting VMs as per earlier email, [[phab:T383583|T383583]] === 2025-01-20 === * 14:23 dhinus: cumin upgraded form 4.2.0 to 5.0.0 on cloudcumin[12]001. patch [[phab:T346453|T346453]] reapplied. === 2025-01-14 === * 21:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T383583|T383583]]) * 21:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T383583|T383583]]) * 20:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T383583|T383583]]) * 20:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T383583|T383583]]) * 20:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T383583|T383583]]) * 20:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T383583|T383583]]) * 18:08 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.roll_restart_osd_daemons (exit_code=0) === 2025-01-13 === * 15:59 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.roll_restart_osd_daemons === 2025-01-10 === * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2025-01-09 === * 20:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 20:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 17:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 17:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 16:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:46 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T309789|T309789]]) * 10:19 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) === 2025-01-08 === * 14:18 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T309789|T309789]]) * 13:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) * 12:59 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T309789|T309789]]) * 12:59 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) * 12:59 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T309789|T309789]]) * 12:59 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) === 2025-01-03 === * 21:01 bd808: `sudo service maintain-dbusers restart` on cloudcontrol1005. Report of missing replica.my.cnf and journalctl output empty due to log rotation. ([[phab:T382962|T382962]]) === 2024-12-18 === * 15:33 arturo: cloudgw failover for [[phab:T382220|T382220]] === 2024-12-04 === * 17:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' * 17:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1035.eqiad.wmnet' * 17:16 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=97) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T380893|T380893]]) * 17:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T380893|T380893]]) * 15:09 andrewbogott: rebooted cloudinfra-cloudvps-puppetserver-1, it's so busy that it's unresponsive * 14:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1035.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T380731|T380731]]) * 14:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1035.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T380731|T380731]]) * 14:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T380893|T380893]]) * 14:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T380893|T380893]]) === 2024-11-27 === * 13:40 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1061.eqiad.wmnet' * 13:26 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' === 2024-11-26 === * 16:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T380731|T380731]]) * 16:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T380731|T380731]]) * 15:16 rook@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:15 rook@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:15 rook@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 15:15 rook@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:14 rook@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/145 * 15:13 rook@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/145 * 15:12 rook@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 15:12 rook@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 13:31 dcaro: added cloudcephmon1004 to the ceph mon pool * 12:41 aborrero@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=255) * 12:40 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 12:35 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 12:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 12:34 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 12:34 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 12:34 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 12:34 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 05:40 andrewbogott: rebooting tools-sgebastion-10.tools.eqiad1.wikimedia.cloud to get NFS things remounted * 04:56 andrewbogott: 'systemctl restart nfs-server' on tools-nfs-2.tools.eqiad1.wikimedia.cloud === 2024-11-25 === * 14:47 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 14:46 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 14:41 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:40 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:55 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 13:55 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 13:48 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:48 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:30 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:29 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:11 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:11 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:28 arturo: create IPv6 networks in eqiad1 ([[phab:T380174|T380174]]), then reverted because network outage * 10:40 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:39 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:33 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 10:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:29 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 10:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-11-23 === * 14:58 taavi: removed broken records for re-created VMs in .eqiad.wmflabs zone === 2024-11-22 === * 14:36 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:35 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:48 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:48 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-11-20 === * 12:56 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:27 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:26 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:13 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 10:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 10:05 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:04 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 10:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2024-11-19 === * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:43 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:43 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 13:37 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:35 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 11:39 arturo: [codfw1dev] performing rabbit full reset [[phab:T380208|T380208]] * 10:26 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 10:24 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 10:24 arturo: [codfw1dev] restart rabbitmq and nova/neutron services for [[phab:T380208|T380208]] === 2024-11-16 === * 08:37 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 08:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 07:52 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 07:52 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-11-15 === * 19:09 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 19:09 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:18 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:18 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:18 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 17:17 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:14 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 17:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:10 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 17:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:06 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 17:06 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 17:06 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 17:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:30 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:29 arturo: [codfw1dev] restart rabbitmq and designate * 16:29 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack === 2024-11-14 === * 09:52 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:48 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack === 2024-11-13 === * 13:05 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/120 * 13:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/120 === 2024-11-11 === * 14:49 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 14:49 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console === 2024-11-08 === * 16:47 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:45 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 16:42 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:42 arturo: [codfw1dev] restart all nova services and rabbitmq out of despair * 16:42 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 13:45 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:44 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 13:33 aborrero@cloudcumin2001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 13:32 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 13:28 arturo: [codfw1dev] restart rabbitmq, openstack services logs show connection errors === 2024-11-07 === * 17:42 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 17:42 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 13:24 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:23 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-11-06 === * 13:09 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:08 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:05 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/117 * 13:04 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/117 * 13:04 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/117 * 13:03 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/117 * 11:15 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:14 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:10 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/116 * 11:09 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/116 * 09:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-11-05 === * 10:24 arturo: [codfw1dev] disable puppet and make changes for testing [[phab:T378192|T378192]] === 2024-11-04 === * 16:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:04 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:04 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:02 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 16:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:57 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:11 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:10 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:09 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:08 aborrero@cloudcumin2001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:08 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 14:08 aborrero@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:08 aborrero@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 13:52 arturo: [codfw1dev] live-hack cloudlb2001-dev and cloudcontrol2004-dev for [[phab:T378192|T378192]] * 13:52 arturo: [codfw1dev] restart rabbitmq * 10:05 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:59 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-29 === * 15:38 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:38 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:40 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:39 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-28 === * 17:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:00 arturo: [codfw1dev] restarting rabbitmq, misbehaving * 17:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:54 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 16:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:53 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 16:53 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:22 dhinus: apt full-upgrade and reboot for cloudcumin* * 16:21 dhinus: upgrade spicerack from 8.8.0 to 8.15.1 on cloudcumin* === 2024-10-24 === * 15:25 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:24 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:51 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:51 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-22 === * 15:02 taavi: recover access to User:Labslogbot [[phab:T376220|T376220]] === 2024-10-21 === * 12:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:06 arturo: [codfw1dev] restart rabbitmq, tofu shows error talking to the designate API === 2024-10-16 === * 14:40 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:39 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-15 === * 15:50 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:50 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:44 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 15:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:38 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:38 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:35 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:35 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:32 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:32 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:31 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:31 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 15:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:41 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:41 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:33 arturo: cloudgw maintenance, firewall change for [[phab:T374714|T374714]] * 10:36 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:36 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:33 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:29 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:28 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:22 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:18 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-11 === * 15:31 arturo: cloudgw maintenance firewall change [[phab:T374716|T374716]] * 09:51 arturo: cloudgw network maintenance related to [[phab:T376879|T376879]] * 08:33 arturo: [codfw1dev] reboot cloudgw2002-dev/2003-dev because network connectivity issues === 2024-10-10 === * 12:04 arturo: manual network failover in cloudgw because maintenance related to [[phab:T376879|T376879]] * 10:40 dhinus: cumin 'cloudrabbit*' 'systemctl restart rabbitmq-server' [[phab:T376802|T376802]] * 09:19 arturo: [codfw1dev] enable IPv6 on cloudgw ([[phab:T374716|T374716]]) === 2024-10-09 === * 14:20 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 14:20 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 14:19 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 14:19 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 14:16 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 14:16 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:58 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:58 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:55 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:55 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:54 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:53 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:53 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:52 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/93 * 10:34 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/95 * 10:34 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/95 * 10:34 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 10:33 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch === 2024-10-08 === * 11:16 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:16 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-07 === * 12:48 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:17 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:06 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:06 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:28 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:11 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:11 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-04 === * 15:12 arturo: cloudservice1005/1006: enable puppet and restore /etc/powerdns/recursor.conf to non-debug mode ([[phab:T374830|T374830]]) * 11:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:39 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:39 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:35 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:32 arturo: cloudservice1005/1006: disable puppet and set `quiet=no` in /etc/powerdns/recursor.conf ([[phab:T374830|T374830]]) === 2024-10-03 === * 14:55 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:54 arturo: [codfw1dev] delete default security group rule list, now tracking them via tofu-infra * 14:53 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 14:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:51 arturo: delete default security group rule list, now tracking them via tofu-infra * 14:48 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 14:47 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-10-02 === * 15:24 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:23 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:23 aborrero@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.tofu (exit_code=97) running tofu plan+apply for main branch * 15:23 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:22 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:21 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:21 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:21 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 12:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:56 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:55 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:43 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:38 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:38 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:26 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch ([[phab:T376211|T376211]]) * 10:21 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch ([[phab:T376211|T376211]]) * 10:16 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch ([[phab:T376211|T376211]]) * 10:15 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch ([[phab:T376211|T376211]]) * 10:05 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch ([[phab:T376211|T376211]]) * 10:04 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch ([[phab:T376211|T376211]]) * 09:14 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/77 ([[phab:T376211|T376211]]) * 09:14 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/77 ([[phab:T376211|T376211]]) * 09:11 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/77 ([[phab:T376211|T376211]]) * 09:10 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/77 ([[phab:T376211|T376211]]) === 2024-10-01 === * 15:41 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 12:36 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:36 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:02 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:53 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:23 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 08:18 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) === 2024-09-30 === * 20:12 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T372814|T372814]]) * 16:59 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.reset_weights (exit_code=0) * 16:41 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 16:35 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T372814|T372814]]) * 13:58 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.reset_weights (exit_code=99) * 13:47 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T372814|T372814]]) * 13:38 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 13:22 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.reset_weights (exit_code=0) * 13:01 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 11:50 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.reset_weights (exit_code=99) * 11:19 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 11:18 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.reset_weights (exit_code=99) * 11:18 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 10:36 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.reset_weights (exit_code=0) * 10:27 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 10:08 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.reset_weights (exit_code=0) * 10:06 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 10:06 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.reset_weights (exit_code=99) * 10:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reset_weights * 10:02 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 10:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 10:00 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=99) * 09:58 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:55 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 09:55 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:49 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 09:48 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 09:48 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:46 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 09:46 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:44 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 09:44 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:41 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 09:40 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:39 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=0) * 09:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:32 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.reset_weights (exit_code=99) * 09:32 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.reset_weights * 09:01 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 09:00 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 09:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) === 2024-09-27 === * 15:27 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:16 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:20 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:20 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:07 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:05 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 13:04 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:02 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:59 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:52 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:51 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:49 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:42 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:41 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 10:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 10:56 arturo: [codfw1dev] restart rabbitmq again * 10:16 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 10:15 arturo: [codfw1dev] restart rabbitmq * 10:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 10:04 arturo: [codfw1dev] enable IPv6 on the neutron virtual router [[phab:T375847|T375847]] === 2024-09-26 === * 14:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:28 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T372814|T372814]]) * 14:28 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T372814|T372814]]) * 10:26 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T372814|T372814]]) * 10:25 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T372814|T372814]]) * 10:25 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 10:11 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:57 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 08:56 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:56 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 08:55 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:55 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 08:55 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:54 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 08:54 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:54 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 08:53 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:52 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 08:04 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 08:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 07:59 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 07:48 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 07:47 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 07:47 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 07:46 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 07:46 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 07:45 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 07:45 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) * 07:42 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T372814|T372814]]) * 07:42 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T372814|T372814]]) === 2024-09-25 === * 14:11 arturo: [codfw1dev] start proxy-02 vm on proxy-codfw1dev project, it was in shutoff mode for unknown reasons * 10:33 arturo: [codfw1dev] cleanup unused security groups ([[phab:T375604|T375604]]) * 09:54 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:53 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:48 arturo: [codfw1dev] restart rabbitmq on all cloudcontrol servers * 09:48 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 09:46 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:46 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:44 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 09:35 arturo: [codfw1dev] deletre a bunch of tests and seemingly unused projects === 2024-09-24 === * 19:33 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 16:00 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 14:36 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T348643|T348643]]) * 14:36 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 14:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:14 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:10 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:56 arturo: [codfw1dev] restart rabbitmq-server on all 3 nodes * 12:53 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:46 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:11 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_rack * 09:10 wmbot~dcaro@urcuchillay: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_rack (exit_code=97) * 09:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_rack === 2024-09-23 === * 19:38 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_rack (exit_code=99) * 15:55 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 15:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:43 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 15:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 15:43 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:42 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:40 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:42 arturo: put cloudvirt1048 in the network-ovs aggregate [[phab:T364457|T364457]] * 14:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_rack * 12:39 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:38 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:37 arturo: [codfw1dev] restart rabbitmq * 12:35 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:30 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:28 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:26 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:24 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:23 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 12:22 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch === 2024-09-21 === * 10:17 dhinus: nova host-evacuate cloudvirt1063 ([[phab:T375223|T375223]]) * 09:59 dhinus: openstack aggregate remove host ceph cloudvirt1063 ([[phab:T375223|T375223]]) * 09:59 dhinus: openstack aggregate add host maintenance cloudvirt1063 ([[phab:T375223|T375223]]) * 09:42 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1063.eqiad.wmnet' * 09:41 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1063.eqiad.wmnet' === 2024-09-20 === * 11:00 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 11:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 10:59 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 10:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 08:27 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T373740|T373740]]) * 08:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T373740|T373740]]) === 2024-09-19 === * 15:46 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T373740|T373740]]) * 15:33 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T373740|T373740]]) * 15:32 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) ([[phab:T373740|T373740]]) * 15:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance ([[phab:T373740|T373740]]) * 10:03 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/51 * 10:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/51 * 09:56 wmbot~dcaro@urcuchillay: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) ([[phab:T374043|T374043]]) * 09:55 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T374043|T374043]]) * 09:51 arturo: [codfw1dev] play with neutron default security group rules (delete, create them, etc) [[phab:T375111|T375111]] * 09:27 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:25 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:25 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:11 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:10 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:08 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:07 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 * 09:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/50 === 2024-09-18 === * 15:37 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/49 * 15:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/49 * 15:24 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/49 * 15:24 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/49 * 15:21 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/49 * 15:21 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/49 * 15:14 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/48 * 15:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/48 * 12:15 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/48 * 12:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/48 * 12:04 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:02 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 12:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/47 * 11:58 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/47 * 11:54 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:53 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:51 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:51 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:50 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:49 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:43 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 11:42 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 11:35 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 11:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 11:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 11:30 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 09:11 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 08:59 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 08:52 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 08:40 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 08:39 dcaro: restarted rabbitmq-server on all cloudrabbits * 08:27 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 08:17 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 08:12 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 08:09 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 08:09 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 08:09 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.openstack.restart_openstack === 2024-09-17 === * 21:24 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T374043|T374043]]) * 16:24 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T374043|T374043]]) * 16:11 wmbot~dcaro@urcuchillay: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) ([[phab:T374043|T374043]]) * 16:11 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T374043|T374043]]) * 15:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 15:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 15:05 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 15:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 15:03 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 15:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 14:32 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 14:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 14:27 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 * 14:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/46 === 2024-09-16 === * 14:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:28 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/45 * 14:28 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/45 * 14:28 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/45 * 14:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/45 * 11:28 arturo: [codfw1dev] created VM bastion-codfw1dev-04 to replace current bastion -03 ([[phab:T374828|T374828]]) === 2024-09-12 === * 10:51 arturo: merging change to keystone wmf hooks https://gerrit.wikimedia.org/r/c/operations/puppet/+/1071230 ([[phab:T374020|T374020]]) === 2024-09-11 === * 16:04 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 16:04 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 16:03 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 16:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 16:01 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 16:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:59 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 15:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 15:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 15:58 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 15:33 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:32 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/44 * 15:30 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/44 * 15:29 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 15:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 15:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:08 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 12:18 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/44 * 12:18 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/44 * 12:15 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 12:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/43 * 12:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/42 * 12:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/42 * 12:08 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/42 * 12:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/42 * 11:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:56 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 11:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:27 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/41 * 11:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/41 * 11:19 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 11:17 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 10:22 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:21 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:20 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 10:19 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:05 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 10:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 07:49 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2004-dev.codfw.wmnet' ([[phab:T374467|T374467]]) * 07:41 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2004-dev.codfw.wmnet' ([[phab:T374467|T374467]]) === 2024-09-10 === * 12:15 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:15 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:14 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:08 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:06 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:05 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:03 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 12:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 11:37 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 11:36 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 10:00 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 10:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:59 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:52 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:42 dcaro: hard-rebooting cloudvirt2004-dev (codfw1dev) having io/hardware issues * 09:22 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:22 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:20 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:19 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:07 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:07 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:03 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 09:00 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 08:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 08:54 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 08:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/40 * 08:46 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 08:45 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) === 2024-09-09 === * 21:32 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T373986|T373986]]) * 16:36 dcaro: cleaned up dns leaks * 15:45 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:44 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:44 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 15:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:24 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/39 * 15:23 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/39 * 15:05 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) * 15:04 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 14:14 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:06 arturo: merged change to cloudgw NAT setting https://gerrit.wikimedia.org/r/c/operations/puppet/+/1071189 * 12:39 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 12:39 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 12:29 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 12:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 11:37 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 11:36 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 11:35 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 11:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:56 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:55 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:54 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:12 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:12 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:11 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 10:11 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:59 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:58 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:58 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:51 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:51 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:49 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:49 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:44 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:44 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:42 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:42 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:40 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:39 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:38 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/38 * 09:15 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) === 2024-09-06 === * 23:18 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T373986|T373986]]) * 18:17 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) * 17:58 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 13:46 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) * 11:13 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 07:27 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) === 2024-09-05 === * 21:32 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 18:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T374043|T374043]]) * 18:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T374043|T374043]]) * 18:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T374043|T374043]]) * 18:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T374043|T374043]]) * 17:32 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) * 17:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T374043|T374043]]) * 17:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T374043|T374043]]) * 17:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T374043|T374043]]) * 16:48 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 16:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T374043|T374043]]) * 16:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T374043|T374043]]) * 16:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T374043|T374043]]) * 14:25 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 14:24 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:19 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 14:17 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 14:16 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/37 * 14:15 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/37 * 13:04 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/37 * 13:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/37 * 13:02 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/37 * 13:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/37 * 12:51 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) * 12:47 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 12:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 10:50 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:49 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:43 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/36 * 10:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/36 * 10:31 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 10:30 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:20 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/35 * 10:19 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/35 * 09:59 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:56 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 09:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:56 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 09:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:55 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/34 * 09:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/34 * 09:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/33 * 09:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/33 * 09:36 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 09:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:34 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/32 * 09:34 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/32 * 09:31 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/32 * 09:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/32 * 09:22 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 09:22 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:20 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/31 * 09:20 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/31 * 09:16 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 09:15 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:14 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 09:12 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 09:10 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 09:10 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 09:03 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 09:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 09:03 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 09:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 08:46 arturo: [codfw1dev] restart rabbitmq @ codfw1dev [[phab:T374002|T374002]] === 2024-09-04 === * 21:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:35 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T373986|T373986]]) * 19:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 17:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 17:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:35 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T373986|T373986]]) * 15:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 15:30 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 15:25 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 15:24 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 15:17 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 15:17 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/30 * 12:18 arturo: [codfw1dev] restart rabbitmq-server.service on all 3 cloudcontrols, all nova-compute agents are down complaining about rabbitmq being unreachable === 2024-09-02 === * 13:27 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 13:27 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 13:23 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 13:23 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 13:19 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 13:19 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 13:07 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 13:06 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:48 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:41 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:40 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:40 aborrero@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.tofu (exit_code=97) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:40 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 11:47 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/29 * 11:46 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/29 * 11:42 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 11:42 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 11:37 dcaro@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:37 dcaro@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:37 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 11:36 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 10:56 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:54 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan+apply for main branch * 10:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:26 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 10:25 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 10:15 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 10:15 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 10:08 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 10:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 === 2024-08-31 === * 13:55 andrewbogott: moving tools-redis-7 off of cloudvirt1048 just in case [[phab:T373740|T373740]] * 13:39 andrewbogott: rebooting cloudvirt1048 from mgmt, it seems to have crashed === 2024-08-30 === * 12:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 11:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-08-29 === * 18:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 18:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2024-08-25 === * 22:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-08-23 === * 14:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T369044|T369044]]) * 14:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T369044|T369044]]) * 00:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 00:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 00:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 00:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-08-22 === * 23:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 23:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 21:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=0) ([[phab:T369044|T369044]]) * 21:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T369044|T369044]]) * 21:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:46 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=97) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T369044|T369044]]) * 20:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.unset_maintenance (exit_code=0) ([[phab:T369044|T369044]]) * 19:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.unset_maintenance ([[phab:T369044|T369044]]) * 19:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:14 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=97) on host 'cloudvirt1039' ([[phab:T369044|T369044]]) * 19:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1039' ([[phab:T369044|T369044]]) * 19:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirt1037' ([[phab:T369044|T369044]]) * 19:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1037' ([[phab:T369044|T369044]]) * 19:14 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirt1035' ([[phab:T369044|T369044]]) * 19:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1035' ([[phab:T369044|T369044]]) * 19:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T369044|T369044]]) * 19:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirt1033' ([[phab:T369044|T369044]]) * 19:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1033' ([[phab:T369044|T369044]]) * 19:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirt1052' ([[phab:T369044|T369044]]) * 19:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1052' ([[phab:T369044|T369044]]) * 18:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 18:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 18:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 18:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 18:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T369044|T369044]]) * 18:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T369044|T369044]]) * 18:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T369044|T369044]]) * 17:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 16:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T369044|T369044]]) * 16:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=0) ([[phab:T369044|T369044]]) * 16:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T369044|T369044]]) === 2024-08-20 === * 03:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) === 2024-08-19 === * 23:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 23:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 23:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 23:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 23:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 23:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 23:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 23:16 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 18:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 18:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2024-08-17 === * 03:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 03:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2024-08-16 === * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 15:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:20 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) * 15:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:10 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 14:40 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 14:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node * 14:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 14:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:33 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:33 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T363344|T363344]]) * 05:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 04:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 04:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 04:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 04:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 04:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 04:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 04:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 04:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 03:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 03:57 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 03:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 03:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 03:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T363344|T363344]]) * 03:50 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 03:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T363344|T363344]]) * 01:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:14 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:12 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:08 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2024-08-15 === * 19:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:20 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:16 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 19:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:33 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=99) * 16:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:27 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 16:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 10:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 09:52 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=0) * 08:56 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 04:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:19 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 04:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 03:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 03:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 03:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 03:03 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 03:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 03:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 03:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 03:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 02:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 02:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:39 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 01:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 01:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 00:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 00:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 00:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 00:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 00:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2024-08-14 === * 23:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 23:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 23:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 23:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 23:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 23:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 23:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 23:22 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:31 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 19:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:06 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 15:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:05 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 15:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 04:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 04:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-08-12 === * 15:44 dcaro@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) ([[phab:T363344|T363344]]) * 11:52 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=99) * 11:51 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T363344|T363344]]) * 11:51 dcaro@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) ([[phab:T363344|T363344]]) * 11:51 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T363344|T363344]]) * 11:51 dcaro@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) ([[phab:T363344|T363344]]) * 11:51 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T363344|T363344]]) * 08:52 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 08:46 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=0) * 08:43 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_rack * 08:37 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=99) * 08:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_rack * 08:37 wmbot~dcaro@urcuchillay: END (ERROR) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=97) ([[phab:T371878|T371878]]) * 08:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_rack ([[phab:T371878|T371878]]) * 08:37 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=99) * 08:33 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_rack * 08:32 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=99) ([[phab:T371878|T371878]]) * 08:27 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_rack ([[phab:T371878|T371878]]) * 08:26 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=99) ([[phab:T371878|T371878]]) * 08:26 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_rack ([[phab:T371878|T371878]]) * 08:17 dcaro@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=97) * 08:17 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_rack === 2024-08-09 === * 19:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T371878|T371878]]) * 18:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T371878|T371878]]) * 13:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 13:38 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) ([[phab:T371878|T371878]]) * 13:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 13:35 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) ([[phab:T371878|T371878]]) * 13:34 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 11:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T371878|T371878]]) * 05:36 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 05:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T371878|T371878]]) * 00:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 00:47 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) ([[phab:T371878|T371878]]) * 00:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) === 2024-08-08 === * 23:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T371878|T371878]]) * 19:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2006-dev.codfw.wmnet' * 18:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' * 18:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2005-dev.codfw.wmnet' * 18:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' * 18:39 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) * 18:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2006-dev.codfw.wmnet' * 18:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 18:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T371878|T371878]]) * 18:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2006-dev.codfw.wmnet' * 18:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 18:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T371878|T371878]]) * 18:25 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 18:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2005-dev.codfw.wmnet' * 18:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T371878|T371878]]) * 18:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2005-dev.codfw.wmnet' * 18:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' * 18:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2004-dev.codfw.wmnet' * 17:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirt2004-dev.codfw.wmnet' * 17:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2004-dev.codfw.wmnet' * 17:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2006-dev.codfw.wmnet' * 17:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2006-dev.codfw.wmnet' * 17:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' * 16:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1001-dev.eqiad.wmnet' * 16:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1001-dev.eqiad.wmnet' * 16:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudbackup1002-dev.eqiad.wmnet' * 16:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1002-dev.eqiad.wmnet' * 16:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudbackup1002-dev.codfw.wmnet' * 16:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudbackup1002-dev.codfw.wmnet' * 16:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2004-dev.codfw.wmnet' * 16:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004-dev.codfw.wmnet' * 16:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2006-dev.codfw.wmnet' * 16:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2006-dev.codfw.wmnet' * 16:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2005-dev.codfw.wmnet' * 16:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:12 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 16:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 16:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' * 15:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2005-dev.codfw.wmnet' * 15:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' * 15:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' * 15:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.wikimedia.org' * 15:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.wikimedia.org' * 15:51 andrewbogott: upgrading codfw1dev to openstack version caracal https://phabricator.wikimedia.org/T369044 * 15:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=0) ([[phab:T369044|T369044]]) * 15:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T369044|T369044]]) * 15:49 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=99) ([[phab:T369044|T369044]]) * 15:49 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T369044|T369044]]) * 15:00 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=99) * 14:51 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 14:01 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=0) * 14:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=0) * 13:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.wait_for_rebalance * 13:00 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=97) * 12:56 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 12:20 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=99) * 12:14 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 11:35 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T371878|T371878]]) * 11:25 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T371878|T371878]]) * 09:43 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 09:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 09:43 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=0) * 07:44 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 07:43 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=99) * 06:41 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 05:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 05:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 00:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 00:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) === 2024-08-07 === * 20:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:11 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=97) * 20:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 20:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 20:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 18:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 18:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 17:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 17:53 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=0) * 17:19 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 17:07 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 17:06 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 17:06 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 17:05 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 17:05 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 17:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 17:03 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 17:03 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 17:03 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 17:02 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 17:02 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 16:56 wmbot~dcaro@urcuchillay: END (ERROR) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=97) * 16:50 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node * 16:50 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:49 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:46 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 16:45 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node * 16:40 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:39 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:38 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:38 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:36 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:35 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:34 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:34 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:33 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:32 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:31 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:31 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:28 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:28 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.undrain_node * 16:25 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:24 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:02 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:02 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 15:41 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:41 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 08:27 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 08:26 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T371878|T371878]]) * 08:11 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T371878|T371878]]) * 06:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T371878|T371878]]) * 03:08 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T371878|T371878]]) * 01:18 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) === 2024-08-06 === * 21:22 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 19:51 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T371878|T371878]]) * 18:36 wmbot~andrew@bullseye: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' * 18:21 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1042.eqiad.wmnet' * 18:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' * 18:17 wmbot~andrew@bullseye: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' * 18:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1043.eqiad.wmnet' * 18:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' * 17:59 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1041.eqiad.wmnet' * 17:58 wmbot~andrew@bullseye: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' * 17:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 17:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' * 17:45 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' * 17:43 wmbot~andrew@bullseye: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' * 17:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:34 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' * 17:32 wmbot~andrew@bullseye: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' * 17:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1045.eqiad.wmnet' * 17:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' * 17:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 17:21 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1038.eqiad.wmnet' * 17:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' * 17:14 wmbot~andrew@bullseye: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' * 17:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047.eqiad.wmnet' * 17:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 17:01 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1037.eqiad.wmnet' * 17:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 17:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 17:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 17:00 wmbot~andrew@bullseye: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' * 17:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:48 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1036.eqiad.wmnet' * 16:47 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1036.eqiad.wmnet' * 16:47 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1036.eqiad.wmnet' * 16:37 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1036.eqiad.wmnet' * 16:37 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1036.eqiad.wmnet' * 16:08 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:04 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:03 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 16:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 16:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 15:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:44 dcaro@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T371878|T371878]]) * 15:43 dcaro@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T371878|T371878]]) * 15:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 15:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 15:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:25 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 15:25 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:22 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1036.eqiad.wmnet' * 15:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1036.eqiad.wmnet' * 15:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 15:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 15:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance === 2024-08-01 === * 13:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 03:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 03:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 03:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 03:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 02:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 02:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 01:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 01:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 01:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-07-31 === * 23:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 23:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 23:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 23:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:04 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 16:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 15:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 15:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 15:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 15:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:27 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:15 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:13 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:13 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:07 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:06 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 12:06 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 11:55 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 11:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/28 * 10:41 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:40 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:37 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:31 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:30 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:29 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:28 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:24 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:24 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:23 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:22 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:21 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:21 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:18 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:17 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:12 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 10:12 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 08:44 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 08:44 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 === 2024-07-30 === * 11:29 arturo: installing nova security updates ([[phab:T371240|T371240]]) === 2024-07-29 === * 18:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:38 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 11:36 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:28 arturo: [codfw1dev] restarting rabbitmq-server on all cloudcontrols, nova-compute cannot contact it * 11:00 arturo: [codfw1dev] installing nova security updates ([[phab:T371240|T371240]]) * 10:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:20 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 08:20 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 === 2024-07-25 === * 14:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 14:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 14:46 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 14:46 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 14:45 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 14:45 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:17 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:16 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:15 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:15 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:12 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:12 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:11 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:11 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:09 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:08 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:08 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:01 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 13:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 12:56 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 12:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 12:55 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 12:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/27 * 11:40 arturo: manually restart maintain-dbusers in cloudcontrol1005 to see if that makes any difference * 11:05 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:03 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:02 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 11:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 11:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 11:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:58 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:57 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:57 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:49 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:48 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 10:47 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/26 * 09:35 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 09:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 08:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 08:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 08:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 08:42 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 08:42 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 08:29 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/25 * 08:28 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/25 * 08:27 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/25 * 08:27 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/25 * 08:25 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/25 * 08:25 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/25 * 08:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 08:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 07:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 07:58 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 === 2024-07-24 === * 16:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 16:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 16:01 aborrero@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.tofu (exit_code=97) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 16:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 15:59 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 15:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 15:57 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:56 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:49 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:49 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:46 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:45 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:43 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:39 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:38 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:37 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:37 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:34 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:34 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:31 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:29 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:28 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 15:28 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/23 * 11:59 arturo: restarted maintain-dbusers.service in cloudcontrol1005, it was stuck doing nothing * 10:45 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 10:45 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 10:44 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 10:44 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 10:43 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 10:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 10:42 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 * 10:42 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/24 === 2024-07-23 === * 13:02 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 13:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 13:00 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 13:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 12:57 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 12:56 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 12:47 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 12:47 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:35 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 10:59 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:59 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:52 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:51 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:51 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:50 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:50 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:48 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:47 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:47 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/22 * 10:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 10:08 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 10:07 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/21 * 10:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/21 * 09:39 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 08:54 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 08:54 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 08:24 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 08:24 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 05:59 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 05:59 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 05:29 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 05:29 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 03:03 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 03:03 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 02:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 02:30 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 00:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 00:04 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) === 2024-07-22 === * 23:34 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 23:33 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 21:07 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 21:07 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 20:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:36 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 18:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:10 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 17:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 17:39 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 16:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/21 * 16:07 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/21 * 16:05 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/21 * 16:05 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/21 * 15:51 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 15:50 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 15:47 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:46 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:46 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:46 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:35 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:31 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 15:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 15:10 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 14:40 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 14:40 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 14:04 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 14:04 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 14:02 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 14:02 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:41 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:41 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:34 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 13:33 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 13:28 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 13:28 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 13:14 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:14 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:11 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:11 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:10 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:06 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:06 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:04 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:04 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:04 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 13:04 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:56 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:55 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:49 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:48 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:47 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:47 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:43 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:35 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:31 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 12:12 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 12:12 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 11:41 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 11:41 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan+apply for main branch * 11:41 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 11:40 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan+apply for main branch * 11:39 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 11:39 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 11:36 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 11:36 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/17 * 11:35 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:35 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:34 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:34 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:31 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.tofu (exit_code=0) running tofu plan for main branch * 11:31 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 11:29 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.tofu (exit_code=99) running tofu plan for main branch * 11:29 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.tofu running tofu plan for main branch * 09:14 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 09:14 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 08:44 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy === 2024-07-21 === * 10:02 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 09:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 09:37 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 09:07 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:07 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 06:40 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 06:40 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 03:45 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 03:45 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 03:15 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 03:14 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 00:50 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 00:50 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 00:20 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 00:20 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) === 2024-07-20 === * 21:56 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 21:56 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 21:26 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 21:26 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 18:59 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:59 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 18:29 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 18:29 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 16:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:04 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:34 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:34 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 13:07 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 13:07 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 12:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 12:37 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 10:11 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 10:11 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 09:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:39 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 07:14 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 07:14 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 06:44 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 06:44 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 04:17 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 04:17 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 03:47 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 03:47 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 01:21 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 01:21 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 00:51 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 00:51 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) === 2024-07-19 === * 22:22 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 22:22 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 21:51 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 21:51 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 19:21 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:21 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 18:51 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 18:51 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 16:25 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:25 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 16:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:53 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:53 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 13:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:26 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 13:26 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 12:55 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 12:55 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 10:26 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 10:26 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 09:56 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:56 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 07:28 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 07:28 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 06:58 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 06:58 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 04:32 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 04:32 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 04:02 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 04:02 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 01:35 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 01:35 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 01:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:04 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) === 2024-07-18 === * 22:38 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 22:38 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 22:22 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 22:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 22:08 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 22:08 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 19:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 19:39 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 19:08 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:08 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 16:40 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 16:40 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 16:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:10 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 13:44 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 13:44 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 13:13 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 13:13 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 12:31 dhinus: upgrade spicerack from 8.5 to 8.8 on cloudcumin* * 10:45 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 10:44 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 10:14 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 10:14 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 07:47 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 07:47 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 07:17 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 07:16 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 04:51 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 04:51 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 04:21 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 04:21 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 03:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 03:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 01:55 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 01:55 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 01:25 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:25 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) === 2024-07-17 === * 22:58 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 22:58 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 22:28 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 22:28 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 20:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:00 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 19:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:29 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 17:02 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 17:02 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 16:32 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:32 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 14:05 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:05 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 13:35 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 13:34 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 11:07 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 11:07 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 10:36 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 10:36 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 08:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 08:10 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 07:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 07:39 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 05:13 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 05:13 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 04:43 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 04:43 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 02:17 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 02:17 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 01:46 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 01:46 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) === 2024-07-16 === * 23:21 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 23:21 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 22:50 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 22:50 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 20:23 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 20:23 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 19:53 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 19:52 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 17:26 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 17:26 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 16:56 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 16:56 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 14:28 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:28 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 13:57 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 13:41 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=0) * 11:20 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 11:19 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 11:14 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 11:02 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 11:02 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 10:57 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 10:25 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 10:21 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.wait_for_rebalance (exit_code=0) * 10:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.wait_for_rebalance * 09:46 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 09:46 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:36 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 09:36 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:36 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 09:31 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 09:19 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 09:19 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:19 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 09:18 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 09:18 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 09:15 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 09:15 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 09:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 08:50 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 08:44 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 08:44 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 08:42 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 08:42 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 08:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 08:35 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 08:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 08:30 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 08:27 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 08:27 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 08:22 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2024-07-15 === * 22:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 22:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:40 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 20:29 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 20:29 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 20:24 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 18:59 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 18:59 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 17:33 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 17:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 17:06 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) * 16:50 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:13 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 14:08 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 14:07 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 14:07 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy === 2024-07-11 === * 13:42 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 13:41 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add * 12:34 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 12:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 11:58 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 11:57 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 11:44 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 11:43 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 02:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 02:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 02:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 02:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 02:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 02:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 02:02 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 02:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:53 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:51 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' * 01:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1061.eqiad.wmnet' * 01:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1061.eqiad.wmnet' === 2024-07-08 === * 17:36 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=0) ([[phab:T309789|T309789]]) * 17:10 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 14:22 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T309789|T309789]]) * 13:01 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) === 2024-07-06 === * 14:06 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-07-05 === * 10:33 arturo: aborrero@cloudcephmon1001:~$ sudo ceph osd unset norebalance * 10:31 arturo: aborrero@cloudcephmon1001:~$ sudo ceph osd unset noin * 08:56 arturo: installing nova/glance/cinder security updates [[phab:T369138|T369138]] === 2024-07-04 === * 20:26 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 20:16 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 20:16 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 14:52 dcaro: rebooting cloudcontrol1007 due to systemd-journal service failing to start * 09:28 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) === 2024-07-03 === * 17:28 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T309789|T309789]]) * 16:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) * 12:28 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 12:22 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) === 2024-07-02 === * 19:17 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 14:24 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) === 2024-07-01 === * 14:28 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T309789|T309789]]) * 14:24 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) * 12:17 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T309789|T309789]]) * 12:03 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) === 2024-06-27 === * 22:50 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 19:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1059.eqiad.wmnet' * 19:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1059.eqiad.wmnet' * 19:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1059.eqiad.wmnet' * 18:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1059.eqiad.wmnet' * 18:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1064.eqiad.wmnet' * 18:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1064.eqiad.wmnet' * 17:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1067.eqiad.wmnet' * 17:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1067.eqiad.wmnet' * 17:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1058.eqiad.wmnet' * 17:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1058.eqiad.wmnet' * 17:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1066.eqiad.wmnet' * 17:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1066.eqiad.wmnet' * 17:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1065.eqiad.wmnet' * 17:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1065.eqiad.wmnet' * 17:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1057.eqiad.wmnet' * 16:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1057.eqiad.wmnet' * 15:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 15:36 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 15:35 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 15:22 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 15:21 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 15:21 wmbot~dcaro@urcuchillay: END (ERROR) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=97) ([[phab:T309789|T309789]]) * 15:21 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 13:44 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T309789|T309789]]) * 12:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirtlocal1001.eqiad.wmnet' * 12:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirtlocal1001.eqiad.wmnet' * 12:07 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) * 12:06 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) * 12:05 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.unset_cluster_maintenance * 12:05 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=0) * 12:04 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 10:32 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) * 10:32 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.unset_cluster_maintenance * 10:31 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=0) * 10:31 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 10:31 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=99) * 10:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 10:30 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=99) * 10:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 10:30 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=99) * 10:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 10:29 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=99) * 10:29 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 10:29 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=99) * 10:28 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 08:36 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=99) * 08:36 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 07:55 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=99) * 07:55 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.set_cluster_in_maintenance === 2024-06-26 === * 22:32 andrewbogott: disabled all g3.* flavors in eqiad1 * 19:18 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 17:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:46 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 15:18 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 15:09 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 14:46 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 14:45 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 14:45 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) ([[phab:T309789|T309789]]) * 14:45 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add ([[phab:T309789|T309789]]) * 11:16 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) ([[phab:T309789|T309789]]) * 10:03 dcaro: taking cloudcephosd1006 out of the pool ([[phab:T348643|T348643]]) * 10:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) * 10:00 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) ([[phab:T309789|T309789]]) * 09:59 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy ([[phab:T309789|T309789]]) === 2024-06-25 === * 22:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2003-dev.codfw.wmnet' * 22:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2003-dev.codfw.wmnet' * 22:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2002-dev.codfw.wmnet' * 22:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2002-dev.codfw.wmnet' * 22:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2001-dev.codfw.wmnet' * 22:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2001-dev.codfw.wmnet' * 22:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt2001-dev.codw.wmnet' * 22:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2001-dev.codw.wmnet' * 21:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2005-dev.codfw.wmnet' * 21:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2005-dev.codfw.wmnet' * 16:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' * 16:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2004-dev.codfw.wmnet' * 16:15 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=97) on host 'cloudvirt2006-dev.codfw.wmnet' * 16:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2006-dev.codfw.wmnet' * 03:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 03:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 03:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 03:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 03:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 03:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-06-24 === * 19:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' * 18:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1056.eqiad.wmnet' * 17:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1055.eqiad.wmnet' * 17:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1055.eqiad.wmnet' * 16:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' * 16:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' === 2024-06-21 === * 10:36 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 10:36 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 10:35 wmbot~arturo@nostromo: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=97) * 10:35 wmbot~arturo@nostromo: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 10:35 wmbot~arturo@nostromo: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 10:35 wmbot~arturo@nostromo: START - Cookbook wmcs.openstack.cloudvirt.vm_console * 09:43 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 09:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 08:31 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T368129|T368129]]) * 08:28 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T368129|T368129]]) * 04:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 4e612eb8-04e1-4541-941d-{{Gerrit|a05519eed60a}} * 04:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 4e612eb8-04e1-4541-941d-{{Gerrit|a05519eed60a}} === 2024-06-20 === * 21:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 14789ac1-bc06-4677-9bb0-{{Gerrit|66c16c887427}} * 21:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 14789ac1-bc06-4677-9bb0-{{Gerrit|66c16c887427}} * 17:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' * 17:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' * 14:38 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 14:38 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 14:08 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' * 14:02 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 14:02 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 14:02 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 14:01 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 13:48 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1052.eqiad.wmnet' * 13:47 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1051.eqiad.wmnet' * 13:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 614f9c99-86f1-410f-8ef5-{{Gerrit|e33d23215ff5}} * 13:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 614f9c99-86f1-410f-8ef5-{{Gerrit|e33d23215ff5}} * 13:25 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1051.eqiad.wmnet' * 13:00 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' * 12:41 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1050.eqiad.wmnet' * 12:40 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1049.eqiad.wmnet' * 12:37 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 12:36 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 12:34 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 12:34 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 11:50 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1049.eqiad.wmnet' * 11:45 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' * 11:13 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1048.eqiad.wmnet' * 11:11 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' * 10:49 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047.eqiad.wmnet' * 10:35 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 10:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 10:34 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 10:34 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 09:38 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' * 09:14 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' * 09:10 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' * 08:55 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1045.eqiad.wmnet' * 08:37 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 08:37 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 04:49 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 77140f83-1a12-43b7-9e47-{{Gerrit|0e779503a525}} * 04:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 77140f83-1a12-43b7-9e47-{{Gerrit|0e779503a525}} * 04:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server d0b1d9d5-1aec-4a05-a2d7-{{Gerrit|6d8522a365dc}} * 04:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server d0b1d9d5-1aec-4a05-a2d7-{{Gerrit|6d8522a365dc}} * 04:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 2d2e8925-b50f-483a-82ee-{{Gerrit|e6a1c588e5be}} * 04:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 2d2e8925-b50f-483a-82ee-{{Gerrit|e6a1c588e5be}} * 04:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 138be95d-93ad-4a85-9245-{{Gerrit|a0a508711555}} * 04:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 270b6533-dc99-4e5d-a642-{{Gerrit|c61138b11891}} * 04:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 270b6533-dc99-4e5d-a642-{{Gerrit|c61138b11891}} * 04:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 138be95d-93ad-4a85-9245-{{Gerrit|a0a508711555}} * 04:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server a3e945dc-3548-47fa-8ce3-{{Gerrit|bf1426ff3b15}} * 04:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 0258b810-29af-448a-af5e-{{Gerrit|ed39e19286df}} * 04:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server a3e945dc-3548-47fa-8ce3-{{Gerrit|bf1426ff3b15}} * 04:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 37e23659-4516-4bd8-a9be-{{Gerrit|4cc55def5560}} * 04:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 0258b810-29af-448a-af5e-{{Gerrit|ed39e19286df}} * 04:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server e94378df-7a48-4b11-b44a-{{Gerrit|bf69aaf132bd}} * 04:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server e94378df-7a48-4b11-b44a-{{Gerrit|bf69aaf132bd}} * 04:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 7da824f3-c9f3-4460-b9b2-{{Gerrit|2a894268a7ca}} * 04:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 37e23659-4516-4bd8-a9be-{{Gerrit|4cc55def5560}} * 04:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 63b82d38-3026-408f-8dcc-{{Gerrit|0ecbd1a3c870}} * 04:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 7da824f3-c9f3-4460-b9b2-{{Gerrit|2a894268a7ca}} * 04:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 7cb371bb-a53a-4e65-a1cf-{{Gerrit|f1a8264a9166}} * 04:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 63b82d38-3026-408f-8dcc-{{Gerrit|0ecbd1a3c870}} * 04:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 6616fcbf-a49e-4e03-b735-{{Gerrit|84d31b535c08}} * 04:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 6616fcbf-a49e-4e03-b735-{{Gerrit|84d31b535c08}} * 04:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 7cb371bb-a53a-4e65-a1cf-{{Gerrit|f1a8264a9166}} * 04:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 6f1171db-9d7d-466f-aaa7-{{Gerrit|cb14e1a6af41}} * 04:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 77140f83-1a12-43b7-9e47-{{Gerrit|0e779503a525}} * 04:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 6f1171db-9d7d-466f-aaa7-{{Gerrit|cb14e1a6af41}} * 04:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 1dc3fcee-cf27-4351-ad60-{{Gerrit|384478624ea3}} * 04:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 77140f83-1a12-43b7-9e47-{{Gerrit|0e779503a525}} * 04:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 1dc3fcee-cf27-4351-ad60-{{Gerrit|384478624ea3}} * 04:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 3df62375-e75d-4068-9c71-{{Gerrit|13519b6bf927}} * 04:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 3af9b40f-d29d-4216-9b64-{{Gerrit|0ebb10f94c7c}} * 04:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 3df62375-e75d-4068-9c71-{{Gerrit|13519b6bf927}} * 04:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 270b6533-dc99-4e5d-a642-{{Gerrit|c61138b11891}} * 04:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 270b6533-dc99-4e5d-a642-{{Gerrit|c61138b11891}} * 04:36 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server ae9fa949-e4ce-4ffb-ad9f-{{Gerrit|4e5a3812d031}} * 04:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server ae9fa949-e4ce-4ffb-ad9f-{{Gerrit|4e5a3812d031}} * 04:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 452dc8d3-6ee0-412c-90ba-{{Gerrit|66f21f9d09c1}} * 04:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 3af9b40f-d29d-4216-9b64-{{Gerrit|0ebb10f94c7c}} * 04:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 452dc8d3-6ee0-412c-90ba-{{Gerrit|66f21f9d09c1}} * 04:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server e94378df-7a48-4b11-b44a-{{Gerrit|bf69aaf132bd}} * 04:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server e94378df-7a48-4b11-b44a-{{Gerrit|bf69aaf132bd}} * 04:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 47b0da1d-e50a-42c1-8cd9-{{Gerrit|dc255cc2f1a3}} * 04:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 47b0da1d-e50a-42c1-8cd9-{{Gerrit|dc255cc2f1a3}} * 02:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T368007|T368007]]) * 02:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T368007|T368007]]) * 02:05 andrewbogott: cloudvirt1063 is unresponsive, cycling power from racadm === 2024-06-19 === * 15:43 taavi@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=97) on host 'cloudvirt1044.eqiad.wmnet' * 15:40 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 15:40 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 15:39 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 15:31 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 15:30 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 15:30 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 15:30 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 15:30 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 15:29 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 12:27 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' * 12:10 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1044.eqiad.wmnet' * 12:10 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' * 11:51 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1043.eqiad.wmnet' * 11:49 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' * 11:33 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1042.eqiad.wmnet' * 01:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server f667d3c2-379a-48d7-ad44-{{Gerrit|4f3933bdb871}} * 01:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server f667d3c2-379a-48d7-ad44-{{Gerrit|4f3933bdb871}} * 01:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 71e4296d-039d-4452-9d92-{{Gerrit|69b9f8eb3aba}} * 01:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 71e4296d-039d-4452-9d92-{{Gerrit|69b9f8eb3aba}} * 01:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server a7725cd2-6162-41a3-8add-{{Gerrit|4dc0668b233b}} * 01:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server a7725cd2-6162-41a3-8add-{{Gerrit|4dc0668b233b}} * 01:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=99) for server 02bb9b5a-cadf-4bee-9b63-{{Gerrit|519b1e9b485b}} * 01:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 02bb9b5a-cadf-4bee-9b63-{{Gerrit|519b1e9b485b}} * 01:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.migrate_server_to_ovs (exit_code=0) for server 02bb9b5a-cadf-4bee-9b63-{{Gerrit|519b1e9b485b}} * 01:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.migrate_server_to_ovs for server 02bb9b5a-cadf-4bee-9b63-{{Gerrit|519b1e9b485b}} === 2024-06-18 === * 21:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.reboot_node (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1006.eqiad.wmnet<nowiki>}</nowiki>' * 20:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.reboot_node on hosts matched by 'D<nowiki>{</nowiki>cloudcontrol1006.eqiad.wmnet<nowiki>}</nowiki>' === 2024-06-17 === * 20:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T364457|T364457]]) * 20:03 andrewbogott: repaced ovs hosts in the 'ceph' aggregate * 19:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T364457|T364457]]) * 19:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T364457|T364457]]) * 19:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T364457|T364457]]) * 19:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T364457|T364457]]) * 19:28 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T364457|T364457]]) * 18:16 andrewbogott: temporarily removing all ovs hosts from the 'ceph' aggregate so the scheduler will stop putting linuxbridge hosts on ovs hosts and breaking them * 17:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T364457|T364457]]) * 17:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T364457|T364457]]) * 17:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T364457|T364457]]) * 17:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T364457|T364457]]) * 13:52 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 13:52 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 12:34 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' * 12:25 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1036.eqiad.wmnet' * 12:03 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 12:02 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 10:53 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' * 10:40 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1035.eqiad.wmnet' === 2024-06-14 === * 14:11 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 14:11 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 13:18 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' * 13:04 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1034.eqiad.wmnet' === 2024-06-13 === * 16:15 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 16:15 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 13:59 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' * 13:40 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1033.eqiad.wmnet' * 13:29 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2003-dev.codfw.wmnet' * 13:24 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2003-dev.codfw.wmnet' === 2024-06-12 === * 11:49 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1031.eqiad.wmnet' * 11:47 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1031.eqiad.wmnet' * 11:37 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1032.eqiad.wmnet' * 11:29 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1032.eqiad.wmnet' * 10:08 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' * 09:50 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1031.eqiad.wmnet' === 2024-06-11 === * 15:55 dcaro: restarting mon service on cloudcephmon1002 to try to release the 2 stuck ops left * 13:30 taavi: pin all existing eqiad1 flavors to linuxbridge hypervisors [[phab:T364458|T364458]] * 12:28 taavi: add all existing eqiad1 cloudvirts to new network-linuxbridge aggregate [[phab:T364458|T364458]] * 10:10 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 10:04 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 10:00 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 09:59 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 09:43 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 09:34 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 09:33 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 09:33 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:33 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-06-07 === * 19:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-06-05 === * 11:57 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.bootstrap_and_add (exit_code=99) * 11:41 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.bootstrap_and_add === 2024-06-04 === * 15:49 taavi: drop hopefully-unused 68.10.in-addr.arpa. from designate [[phab:T361220|T361220]] === 2024-05-30 === * 02:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 02:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-05-29 === * 18:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 18:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 18:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) ([[phab:T364984|T364984]]) * 18:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T364984|T364984]]) * 18:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) ([[phab:T364984|T364984]]) * 18:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T364984|T364984]]) === 2024-05-28 === * 19:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 19:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:28 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 19:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:22 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 19:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 19:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:59 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 18:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:48 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 18:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 14:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-05-21 === * 11:28 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs * 08:18 taavi: stop neutron services on cloudnet1005 [[phab:T364459|T364459]] === 2024-05-20 === * 14:46 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs (exit_code=0) * 14:46 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs * 14:23 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs (exit_code=0) * 14:23 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs * 14:20 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs (exit_code=0) * 14:19 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs * 14:18 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs (exit_code=99) * 14:18 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs * 14:17 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs (exit_code=99) * 14:17 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs * 14:16 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs (exit_code=99) * 14:16 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs * 14:15 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs (exit_code=99) * 14:15 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.migrate_to_ovs === 2024-05-16 === * 08:55 taavi: delete 'monitoring' project https://phabricator.wikimedia.org/T365105 === 2024-05-15 === * 09:08 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T319184|T319184]]) * 08:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T319184|T319184]]) === 2024-05-14 === * 19:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 00:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 00:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 00:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 00:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-05-10 === * 13:23 andrewbogott: deploying updated 2024.1 Horizon === 2024-05-09 === * 19:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-04-25 === * 21:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T356287|T356287]]) * 21:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1032.eqiad.wmnet' ([[phab:T356287|T356287]]) * 21:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T356287|T356287]]) * 20:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt-wdqs1003.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt-wdqs1003.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt-wdqs1002.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt-wdqs1002.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt-wdqs1001.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt-wdqs1001.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:25 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T356287|T356287]]) * 18:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:56 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1006.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) on host 'cloudvirt1033' ([[phab:T356287|T356287]]) * 17:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1033' ([[phab:T356287|T356287]]) * 17:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1006.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1005.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:13 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1005.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T356287|T356287]]) * 17:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1005.eqiad.wmnet' ([[phab:T356287|T356287]]) * 16:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1005.eqiad.wmnet' ([[phab:T356287|T356287]]) * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T356287|T356287]]) * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T356287|T356287]]) * 16:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T356287|T356287]]) * 16:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T356287|T356287]]) * 16:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T356287|T356287]]) * 16:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=99) ([[phab:T356287|T356287]]) * 16:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T356287|T356287]]) === 2024-04-23 === * 09:15 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2002-dev.codfw.wmnet' * 09:10 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2002-dev.codfw.wmnet' === 2024-04-18 === * 12:58 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 12:58 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 12:55 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudvirt2001-dev.codfw.wmnet' * 12:52 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2008-dev.codfw.wmnet' * 12:46 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudvirt2001-dev.codfw.wmnet' * 12:46 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2008-dev.codfw.wmnet' === 2024-04-17 === * 14:12 dcaro: deleting dns leaks * 01:58 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 01:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2006-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 01:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2006-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 01:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2005-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 01:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2005-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 01:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt2004-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 01:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt2004-dev.codfw.wmnet' ([[phab:T356287|T356287]]) === 2024-04-16 === * 21:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2006-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 21:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2006-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 21:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet2005-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 21:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet2005-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 21:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 21:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2005-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 21:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2004-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 20:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 20:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol2001-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2001-dev.codfw.wmnet' ([[phab:T356287|T356287]]) * 19:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2005-dev.codfw.wmnet' ([[phab:T356287|T356287]][A) * 19:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2005-dev.codfw.wmnet' ([[phab:T356287|T356287]][A) * 19:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices2004-dev.codfw.wmnet' ([[phab:T356287|T356287]][A) * 19:20 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004-dev.codfw.wmnet' ([[phab:T356287|T356287]][A) * 19:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2004-dev.codfw.wmnet' ([[phab:T356287|T356287]][A) * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004-dev.codfw.wmnet' ([[phab:T356287|T356287]][A) * 19:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2004.codfw.wmnet' ([[phab:T356287|T356287]]) * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004.codfw.wmnet' ([[phab:T356287|T356287]]) * 19:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudservices2004.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices2004.eqiad.wmnet' ([[phab:T356287|T356287]]) * 19:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=99) ([[phab:T356287|T356287]]) * 19:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T356287|T356287]]) === 2024-04-15 === * 11:19 taavi: update spicerack to 8.5.0 on cloudcumin2001 === 2024-04-10 === * 20:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-04-09 === * 18:13 andrewbogott: rebooting cloudinfra-cloudvps-puppetserver-1; unresponsive * 13:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-04-05 === * 14:37 taavi: run maintain-replica-indexes on all web replicas [[phab:T361945|T361945]] * 14:27 taavi: run maintain-replica-indexes on remaining analytics replicas [[phab:T361945|T361945]] * 14:17 taavi: run maintain-replica-indexes on clouddb1017 [[phab:T361945|T361945]] === 2024-04-04 === * 18:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-04-03 === * 19:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:19 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 19:17 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 19:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:01 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 15:01 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 14:16 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T319184|T319184]]) * 14:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T319184|T319184]]) * 12:40 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 12:40 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 11:48 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T319184|T319184]]) * 11:34 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1039.eqiad.wmnet' ([[phab:T319184|T319184]]) * 11:33 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 11:33 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 10:37 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T319184|T319184]]) * 10:25 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1038.eqiad.wmnet' ([[phab:T319184|T319184]]) * 10:21 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 10:21 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 09:26 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T319184|T319184]]) * 09:17 taavi: manually delete prometheus-node-textfile-wmcs-dnsleaks.service and related files from cloudservices1005/6, leftovers of the designate api to cloudcontrol migration * 09:09 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1037.eqiad.wmnet' ([[phab:T319184|T319184]]) * 08:52 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 08:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance === 2024-04-02 === * 15:01 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=0) * 15:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:00 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 15:00 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 15:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T319184|T319184]]) * 14:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1036.eqiad.wmnet' ([[phab:T319184|T319184]]) * 13:00 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 13:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 12:37 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 12:37 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 12:36 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.depool_and_destroy (exit_code=99) * 12:36 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.depool_and_destroy * 11:45 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T319184|T319184]]) * 11:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1035.eqiad.wmnet' ([[phab:T319184|T319184]]) === 2024-03-27 === * 11:47 taavi: deleting about 2k stale puppet certs by running wmcs-puppetcertleaks in delete mode === 2024-03-22 === * 10:25 dcaro: adding back cloudcephosd1034 to the pool after doing the performance tests ([[phab:T348643|T348643]]) === 2024-03-21 === * 22:07 andrewbogott: doing dist-upgrade on cloudcontrol nodes to get mariadb upgraded for [[phab:T357133|T357133]] * 22:05 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) on host 'cloudcontrol2004-dev.codfw.wmnet' ([[phab:T357133|T357133]]) * 22:04 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol2004-dev.codfw.wmnet' ([[phab:T357133|T357133]]) * 11:08 dcaro: restarting nova-api on cloudcontrol1007 === 2024-03-20 === * 17:02 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 15:39 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 15:27 dcaro: turning off cloudcephosd1030 to swap some disks ([[phab:T348643|T348643]]) * 15:25 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 15:02 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 15:00 wmbot~dcaro@urcuchillay: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T348643|T348643]]) * 14:31 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) === 2024-03-19 === * 18:28 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 15:49 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) * 15:21 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) ([[phab:T348643|T348643]]) * 13:30 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T348643|T348643]]) === 2024-03-18 === * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:22 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 16:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-03-09 === * 16:43 andrewbogott: restarted nova-api on cloudcontrol1006 === 2024-03-08 === * 11:27 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 11:27 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 11:27 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 11:27 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 11:24 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 11:23 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 11:21 arturo: restarted nova-api in cloudcontrol1007, it was complaining about mysql broken pipe * 11:16 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 11:16 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 11:16 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 11:15 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 11:15 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 11:15 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 10:36 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 10:36 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 10:35 taavi@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=97) * 10:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 10:35 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 10:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance === 2024-03-07 === * 12:23 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt2001-dev.codfw.wmnet' * 12:16 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt2001-dev.codfw.wmnet' * 12:16 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) * 12:16 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance === 2024-03-06 === * 17:46 dhinus: running "wmcs-dnsleaks --delete" to clean up 2 leaked records (tools-sgeweblight-10-32) * 15:36 dcaro: renewing puppet ca cert for cloud-puppetmaster-03 * 15:24 dcaro: renewing puppet ca cert for cloudinfra-internal puppetmaster === 2024-03-04 === * 14:56 dhinus: delete project "loggerdiscordbot" in favor of new project "discordbots" [[phab:T358337|T358337]],[[phab:T358427|T358427]] * 12:54 wmbot~dcaro@urcuchillay: END (PASS) - Cookbook wmcs.ceph.reboot_node (exit_code=0) ([[phab:T359049|T359049]]) * 12:48 wmbot~dcaro@urcuchillay: START - Cookbook wmcs.ceph.reboot_node ([[phab:T359049|T359049]]) === 2024-03-01 === * 14:59 taavi: removing wmf-auto-restart-cron from all VMs without cron via cumin - https://gerrit.wikimedia.org/r/c/operations/puppet/+/1007328/ [[phab:T358343|T358343]] * 12:07 dcaro: restarted nova-api on cloudcontrol100* as it was very slow * 12:04 dcaro: restarted nova-api on cloudcontrol1005 as it was very slow === 2024-02-27 === * 18:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-02-26 === * 10:02 arturo: deleting nskaggs account from gerrit's wmcs-trusted group === 2024-02-22 === * 13:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 13:58 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 12:58 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T319184|T319184]]) * 12:57 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T319184|T319184]]) * 12:53 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T319184|T319184]]) * 12:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1034.eqiad.wmnet' ([[phab:T319184|T319184]]) * 11:55 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 11:54 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 09:01 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:00 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-02-21 === * 13:44 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=0) ([[phab:T319184|T319184]]) * 13:43 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance ([[phab:T319184|T319184]]) * 13:20 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T319184|T319184]]) * 13:19 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T319184|T319184]]) * 12:50 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T319184|T319184]]) * 12:50 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T319184|T319184]]) * 12:03 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T319184|T319184]]) * 11:44 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1033.eqiad.wmnet' ([[phab:T319184|T319184]]) === 2024-02-20 === * 11:45 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 11:45 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 11:45 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 11:45 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 11:30 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.post-reimage (exit_code=0) preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) * 11:30 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.post-reimage preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) === 2024-02-19 === * 12:33 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.post-reimage (exit_code=99) preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) * 12:32 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.post-reimage preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) * 12:02 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.post-reimage (exit_code=99) preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) * 12:02 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.post-reimage preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) * 12:00 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.post-reimage (exit_code=99) preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) * 12:00 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.post-reimage preparing cloudvirt cloudvirt1032.eqiad.wmnet for duty (nova discovery, canary VM) Pending aggregates though. ([[phab:T319184|T319184]]) * 10:09 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.pre-reimage (exit_code=0) prepare cloudvirt1032.eqiad.wmnet for reimage (drain, remove nova agent, etc) ([[phab:T319184|T319184]]) * 09:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.pre-reimage prepare cloudvirt1032.eqiad.wmnet for reimage (drain, remove nova agent, etc) ([[phab:T319184|T319184]]) * 09:49 aborrero@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.pre-reimage (exit_code=99) prepare cloudvirt1032.eqiad.wmnet for reimage (drain, remove nova agent, etc) ([[phab:T319184|T319184]]) * 09:49 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.pre-reimage prepare cloudvirt1032.eqiad.wmnet for reimage (drain, remove nova agent, etc) ([[phab:T319184|T319184]]) === 2024-02-15 === * 17:34 wmbot~fran@wmf3169: START - Cookbook wmcs.openstack.roll_reboot_cloudnets ([[phab:T356975|T356975]]) * 15:34 taavi: restart radosgw in eqiad as I am seeing 500 errors * 14:22 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 14:22 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 14:22 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 14:22 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 12:05 aborrero@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T319184|T319184]]) * 11:58 dhinus: restore correct aggregate "localdisk" for cloudvirtlocal1001 and remove "maintenance" * 11:52 aborrero@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T319184|T319184]]) * 05:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' * 05:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1045.eqiad.wmnet<nowiki>}</nowiki>' * 05:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1044.eqiad.wmnet<nowiki>}</nowiki>' * 05:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' * 05:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet<nowiki>}</nowiki>' * 04:52 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1047.eqiad.wmnet<nowiki>}</nowiki>' * 04:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet[B<nowiki>}</nowiki>' * 04:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1046.eqiad.wmnet[B<nowiki>}</nowiki>' * 04:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 04:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' * 04:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 04:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 04:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1049.eqiad.wmnet<nowiki>}</nowiki>' * 04:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1048.eqiad.wmnet<nowiki>}</nowiki>' * 04:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' * 04:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' * 04:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1051.eqiad.wmnet<nowiki>}</nowiki>' * 04:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1050.eqiad.wmnet<nowiki>}</nowiki>' * 04:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' * 04:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' * 03:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1052.eqiad.wmnet<nowiki>}</nowiki>' * 03:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 03:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 03:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1053.eqiad.wmnet<nowiki>}</nowiki>' * 03:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 03:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' * 03:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 03:21 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 03:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1054.eqiad.wmnet<nowiki>}</nowiki>' * 03:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1055.eqiad.wmnet<nowiki>}</nowiki>' === 2024-02-14 === * 22:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' * 21:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' * 21:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1057.eqiad.wmnet<nowiki>}</nowiki>' * 21:38 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1056.eqiad.wmnet<nowiki>}</nowiki>' * 21:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' * 21:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' * 21:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1058.eqiad.wmnet<nowiki>}</nowiki>' * 21:08 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 20:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1059.eqiad.wmnet<nowiki>}</nowiki>' * 20:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' * 20:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 20:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 20:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 20:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1061.eqiad.wmnet<nowiki>}</nowiki>' * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1060.eqiad.wmnet<nowiki>}</nowiki>' * 20:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' * 20:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' * 20:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1062.eqiad.wmnet<nowiki>}</nowiki>' * 20:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' * 20:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 20:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 20:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1064.eqiad.wmnet<nowiki>}</nowiki>' * 20:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1065.eqiad.wmnet<nowiki>}</nowiki>' * 19:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' * 19:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' * 19:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1066.eqiad.wmnet<nowiki>}</nowiki>' * 19:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1067.eqiad.wmnet<nowiki>}</nowiki>' * 19:55 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' * 19:51 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'D<nowiki>{</nowiki>cloudvirtXXXX.eqiad.wmnet<nowiki>}</nowiki>' * 19:51 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirtXXXX.eqiad.wmnet<nowiki>}</nowiki>' * 19:49 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 19:49 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 19:49 wmbot~andrew@bullseye: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) * 19:49 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 19:48 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 19:45 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 19:45 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 19:43 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' * 19:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1063.eqiad.wmnet<nowiki>}</nowiki>' * 19:38 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 19:38 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 19:36 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 19:36 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 19:34 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 19:34 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 19:34 wmbot~andrew@bullseye: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 19:34 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 19:32 wmbot~andrew@bullseye: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 17:13 wmbot~fran@wmf3169: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1001.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T356975|T356975]]) * 17:12 wmbot~fran@wmf3169: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirtlocal1001.eqiad.wmnet<nowiki>}</nowiki>' ([[phab:T356975|T356975]]) * 16:32 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 16:27 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 16:07 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 15:51 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1043.eqiad.wmnet<nowiki>}</nowiki>' * 15:50 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 15:25 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1042.eqiad.wmnet<nowiki>}</nowiki>' * 14:48 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1041.eqiad.wmnet<nowiki>}</nowiki>' * 14:48 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 14:28 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1040.eqiad.wmnet<nowiki>}</nowiki>' * 14:28 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1039.eqiad.wmnet<nowiki>}</nowiki>' * 14:09 taavi: creating some missing $PROJECT.wmcloud.org. DNS zones * 14:07 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1039.eqiad.wmnet<nowiki>}</nowiki>' * 14:06 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1038.eqiad.wmnet<nowiki>}</nowiki>' * 13:50 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:50 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:48 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1038.eqiad.wmnet<nowiki>}</nowiki>' * 13:48 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1037.eqiad.wmnet<nowiki>}</nowiki>' * 13:23 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1037.eqiad.wmnet<nowiki>}</nowiki>' * 13:22 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' * 13:00 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1036.eqiad.wmnet<nowiki>}</nowiki>' * 13:00 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1035.eqiad.wmnet<nowiki>}</nowiki>' * 12:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1035.eqiad.wmnet<nowiki>}</nowiki>' * 12:34 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1034.eqiad.wmnet<nowiki>}</nowiki>' * 12:13 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1034.eqiad.wmnet<nowiki>}</nowiki>' * 12:08 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1033.eqiad.wmnet<nowiki>}</nowiki>' * 11:39 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1033.eqiad.wmnet<nowiki>}</nowiki>' * 11:38 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1032.eqiad.wmnet<nowiki>}</nowiki>' * 11:33 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1032.eqiad.wmnet<nowiki>}</nowiki>' * 11:21 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 10:55 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 10:43 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 10:15 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 09:33 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::virt_ceph<nowiki>}</nowiki>' * 09:31 taavi: failover all dumps traffic to clouddumps1001 * 08:33 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::codfw1dev::virt_ceph<nowiki>}</nowiki>' * 08:16 taavi: reboot clouddumps1001 for kernel updates === 2024-02-13 === * 17:07 wmbot~taavi@runko: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=97) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 17:07 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 16:06 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 16:06 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 16:04 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 16:04 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 16:04 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 16:04 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>O:wmcs::openstack::eqiad1::virt_ceph<nowiki>}</nowiki>' * 16:02 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P<nowiki>{</nowiki>P:openstack::eqiad1::nova::compute::service<nowiki>}</nowiki>' * 16:02 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P<nowiki>{</nowiki>P:openstack::eqiad1::nova::compute::service<nowiki>}</nowiki>' * 16:02 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) on hosts matched by 'P:openstack::eqiad1::nova::compute::service' * 16:02 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'P:openstack::eqiad1::nova::compute::service' * 16:01 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1001.eqiad.wmnet<nowiki>}</nowiki>' * 16:01 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot on hosts matched by 'D<nowiki>{</nowiki>cloudvirt1001.eqiad.wmnet<nowiki>}</nowiki>' * 15:10 wmbot~fran@wmf3169: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) ([[phab:T356975|T356975]]) * 14:55 wmbot~fran@wmf3169: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot ([[phab:T356975|T356975]]) === 2024-02-12 === * 17:33 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=0) * 17:30 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.reboot_node * 17:28 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=0) * 17:25 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.reboot_node * 17:23 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) * 17:21 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudnet.reboot_node * 17:10 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) * 17:06 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 17:04 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) * 16:52 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 16:39 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 16:39 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 16:36 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 16:30 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 16:29 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) * 16:21 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 16:20 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) * 16:14 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 16:04 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) * 15:59 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 15:50 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=0) * 15:46 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 15:43 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) * 15:42 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.reboot_node * 01:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:34 andrewbogott: resetting eqiad1 rabbitmq in hopes of resolving neutron double message warnings * 01:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 00:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 00:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-02-11 === * 21:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 21:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 21:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:57 andrewbogott: running wmcs.openstack.restart_openstack for all eqiad1 services * 20:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 20:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:24 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 11:24 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:23 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 11:23 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-02-08 === * 19:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:43 taavi: deploy change to exclude cloud-private networks from general egress NAT https://phabricator.wikimedia.org/T356850 === 2024-02-06 === * 20:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:31 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 02:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 02:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 02:09 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) * 02:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-02-05 === * 21:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:45 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 20:02 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:55 andrewbogott: rebuilt bookworm base image in eqiad1 with https://gerrit.wikimedia.org/r/c/operations/puppet/+/992677 === 2024-02-02 === * 13:54 arturo: [codfw1dev] cleanup /etc/network/interfaces on cloudlb2003-dev from puppet leftovers === 2024-02-01 === * 10:55 taavi: invite aborrero to /repos/cloud, /toolforge-repos, /cloudvps-repos on gitlab === 2024-01-29 === * 12:54 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:54 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-01-26 === * 09:01 taavi: joining cloudrabbit1001/2 to the cluster on 1003 [[phab:T345610|T345610]] === 2024-01-25 === * 16:48 andrewbogott: taavi just moved all rabbitmq traffic to cloudrabbit1003 as part of [[phab:T345610|T345610]] * 16:45 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:40 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:27 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker-nfs role in the tools cluster * 13:27 wmbot~taavi@runko: Added a new k8s worker-nfs tools-k8s-worker-nfs-2.tools.eqiad1.wikimedia.cloud to the cluster * 13:15 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker-nfs role in the tools cluster * 13:15 wmbot~taavi@runko: Added a new k8s worker-nfs tools-k8s-worker-nfs-1.tools.eqiad1.wikimedia.cloud to the cluster * 12:48 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker-nfs role in the toolsbeta cluster * 12:48 wmbot~taavi@runko: Added a new k8s worker-nfs toolsbeta-test-k8s-worker-nfs-1.toolsbeta.eqiad1.wikimedia.cloud to the cluster === 2024-01-24 === * 11:37 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the toolsbeta cluster * 11:37 taavi@cloudcumin1001: Added a new k8s worker toolsbeta-test-k8s-worker-10.toolsbeta.eqiad1.wikimedia.cloud to the cluster === 2024-01-22 === * 16:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:56 andrewbogott: restarting openstack sevices on eqiad1 to clean up from the mariadb restarts * 15:56 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:07 andrewbogott: merging https://gerrit.wikimedia.org/r/c/operations/puppet/+/992192 and resetting galera cluster in eqiad1 === 2024-01-21 === * 05:19 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 05:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-01-19 === * 23:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 23:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-01-18 === * 20:10 taavi: mysql:labsdbaccounts@m5-master.eqiad.wmnet [labsdbaccounts]> update account_host set status = 'absent' where id = 137613; * 12:38 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 12:38 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-101.tools.eqiad1.wikimedia.cloud to the cluster === 2024-01-17 === * 15:17 andrewbogott: "systemctl restart mariadb@s4.service mariadb@s6.service" on clouddb1015. System is in danger of oom and there are no obvious long queries running === 2024-01-16 === * 14:25 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 14:25 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 14:24 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) for cloudvirt1060.eqiad.wmnet * 14:23 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot for cloudvirt1060.eqiad.wmnet * 14:23 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) for cloudvirt1060.eqiad.wmnet * 14:23 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot for cloudvirt1060.eqiad.wmnet * 13:55 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 13:55 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 13:54 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 13:54 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 13:53 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) * 13:53 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 13:52 wmbot~taavi@runko: END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) * 13:52 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance * 11:35 wmbot~taavi@runko: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) for cloudvirt1060.eqiad.wmnet * 11:34 wmbot~taavi@runko: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot for cloudvirt1060.eqiad.wmnet * 11:22 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) ([[phab:T355061|T355061]]) * 11:02 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot ([[phab:T355061|T355061]]) * 09:41 taavi: drop dbproxy1018/9 grants from all clouddb hosts [[phab:T346947|T346947]] * 09:24 taavi: move cloudvirt2004-dev from 'failed' to 'active' in netbox - seems like that was for [[phab:T348531|T348531]] which is now resolved === 2024-01-15 === * 14:37 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) ([[phab:T355061|T355061]]) * 14:32 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot ([[phab:T355061|T355061]]) * 14:32 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) ([[phab:T355061|T355061]]) * 14:28 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot ([[phab:T355061|T355061]]) * 12:41 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:37 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2024-01-11 === * 10:51 dcaro: restarting striker.service on cloudweb1003 as it seems non-responsive === 2024-01-10 === * 17:43 bd808: Blocking Developer accounts connected to invalid/legacy wikimedia.org email addresses ([[phab:T218239|T218239]]) * 16:48 wmbot~fran@wmf3169: END (PASS) - Cookbook wmcs.do_log_msg (exit_code=0) ([[phab:T346631|T346631]]) * 16:48 wmbot~fran@wmf3169: test message3 from local cookbook ([[phab:T346631|T346631]]) * 16:48 wmbot~fran@wmf3169: START - Cookbook wmcs.do_log_msg ([[phab:T346631|T346631]]) === 2024-01-09 === * 17:50 wmbot~fran@wmf3169: END (PASS) - Cookbook wmcs.do_log_msg (exit_code=0) ([[phab:T346631|T346631]]) * 17:49 wmbot~fran@wmf3169: test message2 from local cookbook ([[phab:T346631|T346631]]) * 17:49 wmbot~fran@wmf3169: START - Cookbook wmcs.do_log_msg ([[phab:T346631|T346631]]) * 17:43 wmbot~fran@wmf3169: %(message)s ([[phab:T346631|T346631]]) * 17:43 wmbot~fran@wmf3169: %(message)s ([[phab:T346631|T346631]]) * 17:42 wmbot~fran@wmf3169: %(message)s ([[phab:T346631|T346631]]) === 2024-01-08 === * 15:52 taavi: verify wmcloud.org, wmflabs.org and toolforge.org in gmail postmaster console to figure out how much google likes us ([[phab:T354112|T354112]]) === 2024-01-07 === * 19:34 andrewbogott: removed cloudvirt1063 from 'ceph' aggregate, added to 'maintenance' aggregate [[phab:T353408|T353408]] * 19:34 andrewbogott: evacuating all VMs from cloudvirt1063. [[phab:T353408|T353408]] === 2024-01-02 === * 16:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) ([[phab:T353408|T353408]]) * 16:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T353408|T353408]]) * 10:22 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=99) * 10:22 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.vm_console === 2023-12-31 === * 21:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:35 andrewbogott: running openstack service restart cookbook in eqiad1 in response to a bunch of service down alerts * 21:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-12-21 === * 16:52 dhinus: puppet node deactivate cloudvirt1063.eqiad.wmnet [[phab:T353406|T353406]] * 03:01 andrewbogott: restarting mariadb on cloudcontrol1005, hoping to get Galera back in sync === 2023-12-20 === * 19:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-12-18 === * 17:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:23 andrewbogott: restarting all eqiad1 openstack services after a rabbitmq upgrade/rebuild for [[phab:T353646|T353646]] * 15:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-12-15 === * 13:13 dcaro: restarted nova-fullstack on codfw as it was stuck (and alerting through stale prometheus file) === 2023-12-14 === * 00:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.set_maintenance (exit_code=99) * 00:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.set_maintenance * 00:25 andrewbogott: evacuating hosts from cloudvirt1063 and depooling. [[phab:T353406|T353406]] === 2023-12-13 === * 16:39 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.toolforge.scale_grid_exec (exit_code=99) === 2023-12-12 === * 21:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:45 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 17:45 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-100.tools.eqiad1.wikimedia.cloud to the cluster * 16:11 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 16:11 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-99.tools.eqiad1.wikimedia.cloud to the cluster * 15:49 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 15:49 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-98.tools.eqiad1.wikimedia.cloud to the cluster === 2023-12-10 === * 18:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:27 andrewbogott: restarting all openstack API servers, hoping to make things a bit more responsive * 18:25 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-12-08 === * 12:00 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.vps.refresh_puppet_certs (exit_code=99) on etcd-discovery-1.cloudinfra-codfw1dev.codfw1dev.wikimedia.cloud ([[phab:T353055|T353055]]) * 11:58 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.vps.refresh_puppet_certs on etcd-discovery-1.cloudinfra-codfw1dev.codfw1dev.wikimedia.cloud ([[phab:T353055|T353055]]) * 11:58 wm-bot2: dcaro@urcuchillay END (ERROR) - Cookbook wmcs.vps.refresh_puppet_certs (exit_code=97) on etcd-discovery-1.cloudinfra-codfw1dev.codfw1dev.wikimedia.cloud * 11:57 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.vps.refresh_puppet_certs on etcd-discovery-1.cloudinfra-codfw1dev.codfw1dev.wikimedia.cloud * 09:38 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) ([[phab:T345084|T345084]]) * 09:32 dcaro: restarting nova and keystone as they are getting too slow ([[phab:T345084|T345084]]) * 09:32 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.openstack.restart_openstack ([[phab:T345084|T345084]]) * 09:32 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) ([[phab:T345084|T345084]]) * 09:31 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.openstack.restart_openstack ([[phab:T345084|T345084]]) === 2023-12-07 === * 13:12 dcaro: rebooting cloudcephosd1001 to make sure puppet7 migration went ok === 2023-12-04 === * 00:08 andrewbogott: rebooting cloudcontrol1006 to recover from full disk error === 2023-12-03 === * 09:05 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:05 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-12-02 === * 12:27 taavi: powercycle cloudvirt1063 [[phab:T352595|T352595]] * 11:28 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 11:28 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-97.tools.eqiad1.wikimedia.cloud to the cluster * 11:02 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 11:02 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-96.tools.eqiad1.wikimedia.cloud to the cluster * 10:50 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 10:50 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-95.tools.eqiad1.wikimedia.cloud to the cluster * 00:21 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 00:21 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-94.tools.eqiad1.wikimedia.cloud to the cluster * 00:18 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 00:18 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-93.tools.eqiad1.wikimedia.cloud to the cluster * 00:15 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 00:15 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-92.tools.eqiad1.wikimedia.cloud to the cluster * 00:06 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 00:06 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-91.tools.eqiad1.wikimedia.cloud to the cluster * 00:05 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 00:05 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-90.tools.eqiad1.wikimedia.cloud to the cluster * 00:01 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=0) for a worker role in the tools cluster * 00:01 taavi@cloudcumin1001: Added a new k8s worker tools-k8s-worker-89.tools.eqiad1.wikimedia.cloud to the cluster === 2023-12-01 === * 17:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:01 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:19 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T351171|T351171]]) * 16:19 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T351171|T351171]]) * 15:49 andrewbogott: reimaging cloudcontrol1005 due to widespread misbehavior * 14:24 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T351171|T351171]]) * 14:20 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T351171|T351171]]) * 14:19 fnegri@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=97) ([[phab:T351171|T351171]]) * 14:18 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T351171|T351171]]) * 13:55 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:54 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:57 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 11:56 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:55 taavi: restart neutron-rpc-server.service on eqiad1 cloudcontrols === 2023-11-30 === * 20:44 andrewbogott: generating to application credentials for the tests that run on tf-infra-test * 19:54 andrewbogott: reimaged cloudrabbit100[23] after https://gerrit.wikimedia.org/r/c/operations/puppet/+/979127. I didn't reimage 1001 because that will require rebuilding the whole cluster but I did remove the related packages. * 18:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1039.eqiad.wmnet' * 18:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1038.eqiad.wmnet' * 18:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1039.eqiad.wmnet' * 18:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1038.eqiad.wmnet' * 18:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1037.eqiad.wmnet' * 18:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1036.eqiad.wmnet' * 18:04 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1035.eqiad.wmnet' * 17:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1037.eqiad.wmnet' * 17:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1036.eqiad.wmnet' * 17:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1035.eqiad.wmnet' * 17:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1033.eqiad.wmnet' * 17:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1034.eqiad.wmnet' * 17:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1032.eqiad.wmnet' * 17:52 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt-wdqs1003.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1034.eqiad.wmnet' * 17:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1033.eqiad.wmnet' * 17:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1032.eqiad.wmnet' * 17:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' * 17:47 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt-wdqs1003.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:47 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt-wdqs1002.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1031.eqiad.wmnet' * 17:42 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt-wdqs1002.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:42 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt-wdqs1001.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:37 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt-wdqs1001.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:36 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:30 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1003.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:30 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:25 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1002.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:25 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:19 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirtlocal1001.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:19 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:14 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1045.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:14 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:09 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1044.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:09 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:04 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1043.eqiad.wmnet' ([[phab:T348843|T348843]]) * 17:04 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:59 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1042.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:59 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:54 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1041.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:54 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:49 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1040.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:45 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:39 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:39 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:34 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:34 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:29 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:21 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:15 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1059.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:15 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:10 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1058.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:10 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:05 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1057.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:05 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T348843|T348843]]) * 16:00 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:56 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:55 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:51 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:51 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:46 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:46 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:41 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:41 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:37 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1051.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:36 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:31 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:26 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:22 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1065.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:22 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:17 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1064.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:17 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:13 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1063.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:12 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:07 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1062.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:07 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:02 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1061.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:02 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:56 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1060.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:49 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:45 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1066.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:43 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:38 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack on host 'cloudvirt1067.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:16 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1006.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:03 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudnet1005.eqiad.wmnet' ([[phab:T348843|T348843]]) * 13:55 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudnet1005.eqiad.wmnet' ([[phab:T348843|T348843]]) * 13:43 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=0) ([[phab:T348843|T348843]]) * 13:43 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudweb.set_maintenance ([[phab:T348843|T348843]]) * 12:18 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T348843|T348843]]) * 12:04 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1005.eqiad.wmnet' ([[phab:T348843|T348843]]) * 11:59 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T348843|T348843]]) * 11:45 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1006.eqiad.wmnet' ([[phab:T348843|T348843]]) * 11:44 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T348843|T348843]]) * 11:28 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudcontrol1007.eqiad.wmnet' ([[phab:T348843|T348843]]) === 2023-11-29 === * 15:27 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) on host 'cloudservices1006.eqiad.wmnet' ([[phab:T348843|T348843]]) * 15:15 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node on host 'cloudservices1006.eqiad.wmnet' ([[phab:T348843|T348843]]) * 14:59 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T348843|T348843]]) * 14:50 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T348843|T348843]]) === 2023-11-28 === * 14:18 taavi: moving wiki replica DNS to use cloudlbs instead of the old proxy VMs [[phab:T346947|T346947]] === 2023-11-27 === * 19:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 19:35 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-24 === * 14:51 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 14:50 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:01 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 12:01 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:01 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 12:00 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:00 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 12:00 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:53 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 11:53 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:53 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 11:53 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:53 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 11:53 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-22 === * 13:28 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 13:21 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 13:02 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 13:00 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-21 === * 10:11 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 10:10 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-20 === * 09:35 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:35 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-17 === * 16:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-16 === * 12:09 taavi@cloudcumin2001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 12:05 taavi@cloudcumin2001: START - Cookbook wmcs.openstack.restart_openstack * 11:23 dhinus: upgraded spicerack from 8.0.2 to 8.0.3 on cloudcumins * 05:51 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 05:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-15 === * 09:50 taavi: move cloudlb hosts to use the nftables firewall backend === 2023-11-14 === * 21:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1056.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1055.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:09 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:06 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 20:00 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1054.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:50 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:21 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1053.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T345811|T345811]]) * 19:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1052.eqiad.wmnet' ([[phab:T345811|T345811]]) * 18:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1050.eqiad.wmnet' ([[phab:T345811|T345811]]) * 18:31 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=97) on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T345811|T345811]]) * 18:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T345811|T345811]]) * 18:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1049.eqiad.wmnet' ([[phab:T345811|T345811]]) * 18:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T345811|T345811]]) * 18:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T345811|T345811]]) * 18:01 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T345811|T345811]]) * 17:59 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:58 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:43 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1048.eqiad.wmnet' ([[phab:T345811|T345811]]) * 17:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T345811|T345811]]) * 17:24 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T345811|T345811]]) * 17:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1047.eqiad.wmnet' ([[phab:T345811|T345811]]) * 12:03 wm-bot2: fran@wmf3169 admin END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T345811|T345811]]) * 11:44 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1046.eqiad.wmnet' ([[phab:T345811|T345811]]) * 10:10 taavi: restart kiwix-mirror-update on clouddumps1001 * 05:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 05:02 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 04:34 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 04:15 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345811|T345811]]) * 04:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 04:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345811|T345811]]) * 03:50 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 03:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 03:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 03:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 03:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 03:26 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 03:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 02:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 02:59 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 02:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 02:59 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 02:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 02:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 02:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 02:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 02:34 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 02:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 02:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 02:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 02:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 02:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 02:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 02:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) === 2023-11-13 === * 22:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345811|T345811]]) * 22:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 22:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 22:11 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 22:09 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 22:09 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 22:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 22:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 21:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 21:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 21:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 21:30 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 21:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 21:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 21:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 21:00 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 19:31 andrewbogott: rebooting cloudcontrol2005-dev, trying to fix general misbehavior * 19:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 19:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) ([[phab:T345811|T345811]]) * 19:16 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 19:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 19:01 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) * 18:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 18:47 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 18:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 18:31 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 18:31 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 18:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 18:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 18:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1033.eqiad.wmnet' * 18:27 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1032.eqiad.wmnet' * 18:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1033.eqiad.wmnet' * 18:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1032.eqiad.wmnet' * 17:09 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=0) ([[phab:T345811|T345811]]) * 17:09 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T345811|T345811]]) * 16:57 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudvirt.unset_maintenance (exit_code=99) ([[phab:T345811|T345811]]) * 16:56 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.unset_maintenance ([[phab:T345811|T345811]]) * 15:37 wm-bot2: fran@wmf3169 admin END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T345811|T345811]]) * 15:19 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1031.eqiad.wmnet' ([[phab:T345811|T345811]]) * 09:11 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 09:08 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 08:56 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 08:51 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-11-11 === * 02:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 02:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 02:41 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 02:14 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 02:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 01:58 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 01:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 01:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 01:21 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 01:05 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 01:03 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) * 00:48 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 00:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 00:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 00:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 00:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 00:46 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 00:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 00:44 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 00:44 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 00:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 00:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 00:40 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) on host 'cloudvirt1058' * 00:39 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1058' * 00:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 00:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.safe_reboot === 2023-11-09 === * 21:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:54 wm-bot2: fran@wmf3169 admin END (PASS) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=0) on host 'cloudvirt1025.eqiad.wmnet' ([[phab:T345811|T345811]]) * 14:52 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.drain on host 'cloudvirt1025.eqiad.wmnet' ([[phab:T345811|T345811]]) * 14:50 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345811|T345811]]) * 14:50 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345811|T345811]]) === 2023-11-08 === * 16:43 andrewbogott: created foundationmemory project for [[phab:T350760|T350760]] === 2023-11-06 === * 20:40 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:36 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:33 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:27 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:25 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 18:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:23 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 18:22 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:20 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 18:19 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 15:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 15:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 15:43 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 15:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 15:42 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) === 2023-11-05 === * 00:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) === 2023-11-04 === * 23:05 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 22:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 20:12 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 16:39 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 16:00 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 13:13 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 04:24 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 01:56 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) === 2023-11-03 === * 16:30 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 16:28 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 16:28 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 16:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 13:13 dhinus: triggering neutron failover from cloudnet1005 to cloudnet1006 ([[phab:T345811|T345811]]) * 06:35 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 01:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 01:46 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 01:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 01:45 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 01:45 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) === 2023-11-02 === * 20:37 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 18:14 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 17:32 taavi: merged cloudcontrol firewall cleanup patch https://gerrit.wikimedia.org/r/c/operations/puppet/+/971211 * 16:46 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) (348643) * 16:23 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 16:23 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 16:21 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 16:00 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 13:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 13:48 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 13:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 11:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 11:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 07:28 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 07:27 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 05:38 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 03:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 03:47 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) (348643) * 03:44 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) === 2023-11-01 === * 23:13 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) (348643) * 21:29 taavi: re-enable puppet on cloudcontrol2006-dev.codfw.wmnet which has fallen off of puppetdb * 21:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 21:15 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 21:15 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 21:14 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 20:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 20:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) (348643) * 19:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) (348643) * 15:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 14:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 14:52 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) (348643) * 14:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 14:50 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 14:49 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 14:49 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 14:48 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 14:47 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 14:47 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 09:04 taavi: reset local cookbook changes on cloudcumin1001 which were causing issues with puppet runs * 09:02 taavi: restart nova-fullstack which had had some issues after yesterday's cloudcontrol1007 reimage * 02:17 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 00:58 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 00:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) === 2023-10-31 === * 23:46 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 18:41 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 14:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 14:11 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 14:10 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 10:16 dhinus: upgrading mariadb-server in cloudcontrol1005 ([[phab:T345811|T345811]]) * 10:04 dhinus: upgrading mariadb-server in cloudcontrol1006 ([[phab:T345811|T345811]]) * 09:51 dhinus: upgrading mariadb-server in cloudcontrol1007, second attempt ([[phab:T345811|T345811]]) * 06:16 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) (348643) * 05:20 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 02:27 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) (348643) * 02:27 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node (348643) * 02:27 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.ceph.osd.drain_node (exit_code=97) (348643) * 02:26 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 01:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 01:59 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 01:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 01:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) * 01:07 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) (348643) === 2023-10-30 === * 17:40 andrewbogott: rebooting tools-db-1.tools.eqiad1.wikimedia.cloud for yet another oom death * 17:09 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 17:04 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 17:02 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 16:56 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) (348643) * 16:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) (348643) * 16:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 16:55 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.drain_node (348643) * 16:54 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:53 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:53 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:52 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:51 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:39 dhinus: upgrading mariadb-server in cloudcontrol1007 ([[phab:T345811|T345811]]) * 16:38 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T348643|T348643]]) * 16:38 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T348643|T348643]]) * 16:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T348643|T348643]]) * 16:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T348643|T348643]]) * 16:37 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T348643|T348643]]) * 16:37 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T348643|T348643]]) * 16:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T348643|T348643]]) * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T348643|T348643]]) * 16:30 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T348643|T348643]]) * 16:30 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T348643|T348643]]) * 16:29 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) ([[phab:T348643|T348643]]) * 16:29 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node ([[phab:T348643|T348643]]) * 16:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node * 16:17 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 16:17 andrew@cloudcumin1001: START - Cookbook wmcs.ceph.osd.undrain_node === 2023-10-28 === * 07:06 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) === 2023-10-27 === * 15:38 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 15:38 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 15:31 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 15:29 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 15:29 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 15:28 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 15:27 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 15:26 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 13:21 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 13:09 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 12:26 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 12:13 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 12:11 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 10:10 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 10:09 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 09:05 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 09:04 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node === 2023-10-26 === * 14:02 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 09:46 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node === 2023-10-25 === * 11:09 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=99) for a ingress role in the toolsbeta cluster * 11:00 taavi: update cloudcumins to spicerack 8.x * 10:34 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.toolforge.add_k8s_node (exit_code=99) for a ingress role in the toolsbeta cluster === 2023-10-24 === * 15:31 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:30 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-10-23 === * 16:48 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 16:46 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:27 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:27 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:24 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 15:22 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:06 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.vps.create_project (exit_code=99) for project catalyst in eqiad1 * 15:06 wm-bot2: fran@wmf3169 START - Cookbook wmcs.vps.create_project for project catalyst in eqiad1 * 10:36 taavi: merged change https://gerrit.wikimedia.org/r/c/operations/puppet/+/966494 which touches the pdns web server config === 2023-10-20 === * 15:17 dcaro: upgraded cloudcephosd1004 to v15 ([[phab:T349363|T349363]]) * 14:25 dcaro: upgraded cloudcephosd1003 to v15 ([[phab:T349363|T349363]]) * 13:20 dcaro: upgraded cloudcephosd1002 to v15 ([[phab:T349363|T349363]]) * 10:33 dcaro: upgraded cloudcephosd1001 to v15 ([[phab:T349363|T349363]]) * 08:26 dcaro: codfw ceph enabled diskprediction_local module, will take a bit to populate/start getting predictions ([[phab:T348716|T348716]]) === 2023-10-17 === * 15:29 taavi@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T349109|T349109]]) * 15:28 taavi@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T349109|T349109]]) === 2023-10-16 === * 03:32 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 03:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-10-13 === * 23:12 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 23:10 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 19:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 19:24 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 17:05 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 17:03 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:43 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:42 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:36 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:33 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:18 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:15 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 08:41 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 08:31 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 08:30 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 08:20 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) === 2023-10-12 === * 17:16 wm-bot2: fran@wmf3169 END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) ([[phab:T341285|T341285]]) * 17:16 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) === 2023-10-11 === * 10:42 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) ([[phab:T341285|T341285]]) * 10:36 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack ([[phab:T341285|T341285]]) * 10:13 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 07:03 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 06:41 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node === 2023-10-10 === * 17:25 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 14:50 dcaro: removing ~100 dangling backup snapshots from eqiad1-compute ceph pool * 12:46 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) ([[phab:T341285|T341285]]) * 12:41 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack ([[phab:T341285|T341285]]) * 12:23 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 11:57 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 11:38 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) ([[phab:T341285|T341285]]) * 11:33 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack ([[phab:T341285|T341285]]) * 11:33 wm-bot2: fran@wmf3169 admin END (FAIL) - Cookbook wmcs.openstack.cloudvirt.safe_reboot (exit_code=99) * 11:32 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.safe_reboot * 11:00 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=99) ([[phab:T341285|T341285]]) * 11:00 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack ([[phab:T341285|T341285]]) * 10:59 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) ([[phab:T341285|T341285]]) * 10:52 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack ([[phab:T341285|T341285]]) * 10:03 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) ([[phab:T341285|T341285]]) * 09:56 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack ([[phab:T341285|T341285]]) * 09:50 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack (exit_code=0) ([[phab:T341285|T341285]]) * 09:43 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.live_upgrade_openstack ([[phab:T341285|T341285]]) * 08:19 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 08:18 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 08:18 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 08:17 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 08:17 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 08:17 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 08:17 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 08:13 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 08:13 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node === 2023-10-09 === * 17:16 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 16:26 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 16:18 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 15:49 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) ([[phab:T341285|T341285]]) * 15:41 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 15:26 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) ([[phab:T341285|T341285]]) * 13:55 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 13:45 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 13:38 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 13:13 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 13:12 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 13:04 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 12:33 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 09:10 dcaro: undrained cephosd1011 * 09:09 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 09:09 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 09:06 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 09:05 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 09:05 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 09:05 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 07:35 taavi: restart postgresql on cloudbackup2001 [[phab:T348431|T348431]] === 2023-10-05 === * 16:41 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 16:29 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 16:24 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 16:11 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 15:30 arturo: operating on cloudgw @ eqiad1 ([[phab:T347469|T347469]]) * 14:07 wm-bot2: dcaro@urcuchillay admin END (FAIL) - Cookbook wmcs.ceph.osd.drain_rack (exit_code=99) * 12:55 arturo: doing cloudgw maintenance operations [[phab:T347469|T347469]] * 11:54 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_rack * 10:57 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 09:54 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 09:52 arturo: [codfw1dev] aborrero@cloudcontrol2001-dev:~ $ sudo keystone-manage fernet_setup --keystone-user keystone --keystone-group keystone * 09:47 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 09:40 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 09:33 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 08:50 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 07:49 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 07:32 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 00:04 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) === 2023-10-04 === * 20:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 20:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:52 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 15:55 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 14:54 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=0) ([[phab:T341285|T341285]]) * 14:44 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 14:41 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) ([[phab:T341285|T341285]]) * 14:40 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node ([[phab:T341285|T341285]]) * 13:50 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) ([[phab:T341285|T341285]]) * 12:09 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 12:08 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 07:21 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 07:21 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 07:15 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 07:12 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 07:11 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node === 2023-10-03 === * 19:38 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 13:38 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 12:35 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 12:33 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 12:32 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 12:32 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 12:29 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 12:01 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 11:33 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 08:58 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 08:43 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.drain_node (exit_code=0) * 08:42 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 08:39 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 08:38 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 08:23 dcaro: set .rgw.root pool on eqiad as rgw app (`ceph osd pool application enable .rgw.root rgw`) === 2023-10-02 === * 17:39 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 16:15 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 16:13 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 14:30 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 13:58 taavi: cloudcontrol1005,7: `sudo systemctl reset-failed keystone_sync_keys_from_cloudcontrol1006.eqiad.wmnet.service` * 13:52 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 13:40 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) * 13:38 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node * 13:37 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T316544|T316544]]) * 12:26 arturo: [codfw1dev] run `update domains set master = '185.15.57.25:5354 185.15.57.26:5354 172.20.5.9:5354 172.20.5.8:5354';` in cloudservies2005-dev * 11:55 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T316544|T316544]]) * 11:55 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T316544|T316544]]) * 11:55 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T316544|T316544]]) * 11:43 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) ([[phab:T316544|T316544]]) * 08:44 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.drain_node ([[phab:T316544|T316544]]) === 2023-09-29 === * 12:43 taavi: taavi@cloudcontrol1005 ~ $ os subnet set a69bdfad-d7d2-4cfa-8231-{{Gerrit|3d6d3e0074c9}} --no-dns-nameservers --dns-nameserver 172.20.255.1 * 08:36 taavi: start script to fix networking on broken bullseye instances [[phab:T347665|T347665]] === 2023-09-28 === * 20:29 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 20:26 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 18:44 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 18:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:01 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=0) * 12:01 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 12:01 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.osd.undrain_node (exit_code=99) * 12:01 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.osd.undrain_node * 09:48 arturo: rebooting cloudgw1001/1002 for sysctl and kernel upgrades === 2023-09-27 === * 12:07 arturo: merging cloudgw firewall changes https://gerrit.wikimedia.org/r/c/operations/puppet/+/961360 * 09:36 taavi: move maintain-dbusers to cloudcontrol1005 * 01:34 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:29 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-26 === * 17:07 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 17:06 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-24 === * 15:39 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 15:37 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:35 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 15:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-23 === * 08:40 taavi: restart keystone === 2023-09-22 === * 14:03 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 14:02 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 14:02 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.drain_node (exit_code=0) * 14:01 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 14:01 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.undrain_node (exit_code=0) * 14:00 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.undrain_node * 14:00 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 13:57 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 13:56 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 13:55 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 13:53 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.undrain_node (exit_code=0) * 13:52 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.undrain_node * 13:49 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.undrain_node (exit_code=99) * 13:48 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.undrain_node * 13:48 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.undrain_node (exit_code=99) * 13:47 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.undrain_node * 13:47 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.drain_node (exit_code=0) * 13:46 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 13:46 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 13:44 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 13:43 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 13:43 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 13:41 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 13:41 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 13:04 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.drain_node (exit_code=0) * 13:03 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 12:37 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 12:37 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 12:33 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 12:33 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 12:33 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 12:28 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node * 11:43 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.ceph.drain_node (exit_code=99) * 11:42 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.drain_node === 2023-09-21 === * 16:35 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.refresh_puppet_certs (exit_code=0) on tools-db-3.tools.eqiad1.wikimedia.cloud * 16:34 fnegri@cloudcumin1001: START - Cookbook wmcs.vps.refresh_puppet_certs on tools-db-3.tools.eqiad1.wikimedia.cloud * 02:12 andrew@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.restart_openstack (exit_code=97) * 02:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-20 === * 21:49 root@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:45 root@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 21:38 root@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 21:35 root@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 21:34 root@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 21:31 root@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:26 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 16:23 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 11:12 arturo: moving openstack API endpoint to cloudlb ([[phab:T346439|T346439]]) * 10:28 arturo: running SQL command `update domains set master="172.20.1.5:5354 172.20.2.4:5354 185.15.56.162:5354 185.15.56.163:5354";` on cloudservices1005/1006 ([[phab:T346042|T346042]]) * 10:06 arturo: running SQL command `update domains set master="185.15.56.162:5354 185.15.56.163:5354"` on cloudservices1005/1005 ([[phab:T346042|T346042]]) === 2023-09-19 === * 18:54 andrewbogott: depooling clouddb1019 to let it recover from high memory use. [[phab:T346826|T346826]] * 15:57 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 15:53 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:47 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 15:46 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-18 === * 11:53 taavi: update designate urls in Keystone to point to openstack-next, until cloudlb is serving the main openstack address [[phab:T346042|T346042]] * 11:40 arturo: decomission cloudservices1005 [[phab:T346042|T346042]] in preparation for re-racking * 08:45 arturo: hardcode `185.15.56.161 openstack.eqiad1.wikimediacloud.org` in /etc/hosts in cloudcontrol1005 for [[phab:T346441|T346441]] === 2023-09-15 === * 11:43 arturo: merging NAT change for [[phab:T346426|T346426]] in cloudgw * 10:33 arturo: faiolver cloudgw1001 into cloudgw1002, investigating a nftables syntax error ([[phab:T346432|T346432]]) === 2023-09-14 === * 21:25 wm-bot2: andrew@bullseye END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 21:25 wm-bot2: andrew@bullseye START - Cookbook wmcs.openstack.cloudvirt.drain * 21:23 wm-bot2: andrew@bullseye END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 21:23 wm-bot2: andrew@bullseye START - Cookbook wmcs.openstack.cloudvirt.drain * 21:18 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) * 21:18 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain * 17:13 wm-bot2: fran@wmf3169 admin END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345810|T345810]]) * 17:06 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345810|T345810]]) * 17:01 wm-bot2: fran@wmf3169 admin END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345810|T345810]]) * 16:56 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345810|T345810]]) * 16:51 wm-bot2: fran@wmf3169 admin END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345810|T345810]]) * 16:51 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345810|T345810]]) * 16:29 wm-bot2: fran@wmf3169 admin END (FAIL) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=99) ([[phab:T345810|T345810]]) * 16:29 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345810|T345810]]) * 16:25 fnegri@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudvirt.drain (exit_code=97) ([[phab:T345810|T345810]]) * 16:25 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudvirt.drain ([[phab:T345810|T345810]]) * 16:12 topranks: DNS operation: remove old DNS entry for ns0.openstack.eqiad1.wikimediacloud.org. on wikimedia authdns (was pointing to 208.80.154.148) * 16:10 topranks: DNS operation: add new DNS entry for ns0.openstack.eqiad1.wikimediacloud.org. on wikimedia authdns pointing to 185.15.56.162 * 14:42 arturo: DNS operation: route 208.80.154.148 to cloudservices1006 in anticipation of cloudservices1005 decom ([[phab:T346042|T346042]]) * 12:11 arturo: enable puppet on cloudservices1006 to drop local NAT hacks and enable new DNS auth IP address ([[phab:T346042|T346042]]) === 2023-09-13 === * 17:11 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=0) ([[phab:T345811|T345811]]) * 17:08 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudnet.reboot_node ([[phab:T345811|T345811]]) * 17:04 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) ([[phab:T345811|T345811]]) * 17:04 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudnet.reboot_node ([[phab:T345811|T345811]]) * 16:57 wm-bot2: dcaro@urcuchillay END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) ([[phab:T345811|T345811]]) * 16:57 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.openstack.cloudnet.reboot_node ([[phab:T345811|T345811]]) * 16:56 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) ([[phab:T345811|T345811]]) * 16:56 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudnet.reboot_node ([[phab:T345811|T345811]]) * 16:53 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) ([[phab:T345811|T345811]]) * 16:53 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudnet.reboot_node ([[phab:T345811|T345811]]) * 16:49 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) ([[phab:T345811|T345811]]) * 16:49 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudnet.reboot_node ([[phab:T345811|T345811]]) * 16:41 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudnet.reboot_node (exit_code=99) ([[phab:T345811|T345811]]) * 16:40 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudnet.reboot_node ([[phab:T345811|T345811]]) * 01:57 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 01:57 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 01:55 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 01:55 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 01:54 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 01:54 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-12 === * 18:06 andrewbogott: update domains set master='185.15.56.163:5354 208.80.154.11:5354 10.64.151.4:5354'; on cloudservices1005 + cloudservices1006 * 17:59 andrewbogott: mysql:root@localhost [pdns]> update domains set master='185.15.56.163:5354 208.80.154.11:5354'; * 17:59 andrewbogott: "designate-manage pool update' on cloudservices1005 to remove cloudservices1004 from the pool * 17:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 17:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:16 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:12 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 15:11 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 15:07 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-11 === * 12:36 arturo: update DNS resolver cloud-wide to use 172.20.255.1 ([[phab:T342621|T342621]]) === 2023-09-10 === * 02:52 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 02:49 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 02:42 andrew@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) * 02:41 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-09-07 === * 10:09 dhinus: reimaging cloudcontrol2001-dev to bookworm ([[phab:T345810|T345810]]) === 2023-09-06 === * 16:56 fnegri@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) * 16:56 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:53 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:52 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:51 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:51 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:51 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:51 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:51 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:51 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:37 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:37 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:37 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:35 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:35 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:34 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=99) * 16:34 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:23 fnegri@cloudcumin1001: END (ERROR) - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node (exit_code=97) * 16:23 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudcontrol.upgrade_openstack_node * 16:13 fnegri@cloudcumin1001: END (FAIL) - Cookbook wmcs.openstack.cloudweb.set_maintenance (exit_code=99) * 16:12 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.cloudweb.set_maintenance === 2023-09-05 === * 12:34 arturo: synced pdns database from cloudservices1004 to cloudservices1006 ([[phab:T345240|T345240]]) * 12:18 arturo: updating pools.yaml in all cloudservices designate nodes ([[phab:T345240|T345240]]) * 10:54 arturo: running SQL command `update domains set master="208.80.154.11:5354 208.80.154.148:5354 10.64.151.4:5354";` on all 3 cloudservices nodes ([[phab:T345240|T345240]]) === 2023-09-04 === * 15:58 arturo: stop and mask designate-sink.service @ cloudservices1006 * 14:19 arturo: started all designate services on cloudservices1006 [[phab:T345240|T345240]] * 10:46 arturo: added designate galera DB grants for cloudlb [[phab:T345240|T345240]] * 08:40 arturo: stopped all designate services on cloudservices1006 [[phab:T345240|T345240]] === 2023-09-01 === * 10:28 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.openstack.network.tests (exit_code=1) ([[phab:T345282|T345282]]) * 10:25 wm-bot2: fran@wmf3169 START - Cookbook wmcs.openstack.network.tests ([[phab:T345282|T345282]]) === 2023-08-31 === * 15:14 wm-bot2: fran@wmf3169 END (FAIL) - Cookbook wmcs.toolforge.grid.get_cluster_status (exit_code=99) * 15:14 wm-bot2: fran@wmf3169 START - Cookbook wmcs.toolforge.grid.get_cluster_status * 12:49 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.openstack.cloudnet.show (exit_code=0) * 12:49 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.openstack.cloudnet.show * 12:48 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) * 12:48 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.unset_cluster_maintenance * 12:47 wm-bot2: dcaro@urcuchillay END (PASS) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=0) * 12:47 wm-bot2: dcaro@urcuchillay START - Cookbook wmcs.ceph.set_cluster_in_maintenance * 12:46 wm-bot2: fran@wmf3169 END (PASS) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=0) * 12:46 wm-bot2: fran@wmf3169 START - Cookbook wmcs.ceph.set_cluster_in_maintenance === 2023-08-30 === * 13:07 wm-bot2: dcaro testing stuff === 2023-08-28 === * 15:05 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:05 wm-bot2: Restarting openstack services on cloudservices1005: ['designate-producer', 'designate-sink', 'designate-worker', 'designate-central', 'designate-mdns', 'designate-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudservices1004: ['designate-worker', 'designate-api', 'designate-mdns', 'designate-producer', 'designate-central', 'designate-sink', 'designate-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:05 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:04 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:03 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:02 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:01 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:01 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:01 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] ([[phab:T345084|T345084]]) - cookbook ran by root@cloudcumin1001 * 15:01 fnegri@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-08-23 === * 16:10 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 16:10 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by root@cloudcumin1001 * 16:10 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by root@cloudcumin1001 * 16:10 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 16:10 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 16:09 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by root@cloudcumin1001 * 16:09 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by root@cloudcumin1001 * 16:09 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 16:09 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 16:09 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 16:09 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 16:08 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 16:08 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 16:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-08-15 === * 19:33 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 19:33 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 19:33 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 19:32 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 19:32 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 19:32 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 16:01 andrewbogott: rebooting cloudvirt2001-dev in an attempt to figure out what's happening with bastions * 15:42 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 15:42 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by root@cloudcumin1001 * 15:42 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by root@cloudcumin1001 * 15:42 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 15:42 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 15:42 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by root@cloudcumin1001 * 15:42 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by root@cloudcumin1001 * 15:41 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 15:41 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 15:41 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 15:40 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 15:40 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 15:40 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 15:40 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack * 12:39 dcaro: removed some logs from the cloudmetrics1003:/var/log/carbon/ directory and stopped the carbon processes (they were crashing and filling up the disk with logs) === 2023-08-14 === * 22:11 andrew@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) * 22:11 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by root@cloudcumin1001 * 22:11 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by root@cloudcumin1001 * 22:11 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 22:11 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 22:10 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by root@cloudcumin1001 * 22:10 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by root@cloudcumin1001 * 22:10 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 22:09 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 22:09 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 22:09 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by root@cloudcumin1001 * 22:09 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 22:08 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by root@cloudcumin1001 * 22:08 andrew@cloudcumin1001: START - Cookbook wmcs.openstack.restart_openstack === 2023-08-07 === * 09:39 taavi: cloud vps graphite service was disabled: https://wikitech.wikimedia.org/wiki/News/2023_Cloud_VPS_metrics_changes [[phab:T326266|T326266]] === 2023-08-03 === * 13:07 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.do_log_msg (exit_code=0) * 13:07 fnegri@cloudcumin1001: START - Cookbook wmcs.do_log_msg * 13:07 wm-bot2: Test SAL log ([[phab:T341793|T341793]]) - cookbook ran by root@cloudcumin1001 === 2023-07-31 === * 16:16 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.do_log_msg (exit_code=0) * 16:16 wm-bot2: Test SAL log ([[phab:T325756|T325756]]) - cookbook ran by root@cloudcumin1001 * 16:15 fnegri@cloudcumin1001: START - Cookbook wmcs.do_log_msg * 15:33 wm-bot2: Test SAL log ([[phab:T325756|T325756]]) - cookbook ran by root@cloudcumin1001 * 14:42 wm-bot2: Test SAL log ([[phab:T325756|T325756]]) - cookbook ran by root@cloudcumin1001 * 14:35 wm-bot2: Test SAL log ([[phab:T325756|T325756]]) - cookbook ran by root@cloudcumin1001 * 13:55 andrewbogott: recreating the codfw1dev galera cluster according to https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Galera -- mariadb is stopped (and won't start) on all three cloudcontrol nodes * 13:39 wm-bot2: Test SAL log ([[phab:T325756|T325756]]) - cookbook ran by root@cloudcumin1001 === 2023-07-27 === * 10:52 arturo: adding cloud-private subnet to cloudnet1005/1006 hosts in eqiad1 ([[phab:T342619|T342619]]) === 2023-07-19 === * 14:02 wm-bot2: Draining cloudvirt2002-dev.codfw.wmnet ([[phab:T335840|T335840]]) - cookbook ran by raymond@ubuntu * 14:02 wm-bot2: Safe rebooting cloudvirt2002-dev.codfw.wmnet ([[phab:T335840|T335840]]) - cookbook ran by raymond@ubuntu === 2023-07-17 === * 15:55 arturo: cloudcontrol1005 was shutdown earlier today ([[phab:T341495|T341495]]) * 15:35 arturo: [codfw1dev] cloudweb2002-dev up and running after reracking ([[phab:T327919|T327919]]) * 15:11 arturo: [codfw1dev] powered off cloudweb2002-dev for reracking ([[phab:T327919|T327919]]) * 12:45 taavi: removing diamond from remaining buster instances [[phab:T317032|T317032]] === 2023-07-13 === * 10:48 wm-bot2: Restarting openstack services on cloudservices1005: ['designate-producer', 'designate-sink', 'designate-worker', 'designate-central', 'designate-mdns', 'designate-agent'] - cookbook ran by dcaro@urcuchillay * 10:48 wm-bot2: Restarting openstack services on cloudservices1004: ['designate-worker', 'designate-api', 'designate-mdns', 'designate-producer', 'designate-central', 'designate-sink', 'designate-agent'] - cookbook ran by dcaro@urcuchillay * 10:48 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] - cookbook ran by dcaro@urcuchillay * 10:48 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] - cookbook ran by dcaro@urcuchillay * 10:48 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:47 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:46 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:45 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by dcaro@urcuchillay * 10:45 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by dcaro@urcuchillay * 10:45 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:45 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:45 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:45 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:44 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:43 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:43 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:43 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:43 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:43 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:43 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:42 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by dcaro@urcuchillay * 10:42 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:42 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:42 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:42 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:42 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:42 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:38 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay * 10:33 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by dcaro@urcuchillay === 2023-07-07 === * 10:28 taavi: backfilling <nowiki>{</nowiki>project<nowiki>}</nowiki>.wmcloud.org and other currently-named DNS zones to projects that don't have them === 2023-07-06 === * 17:02 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:00 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 17:00 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-07-04 === * 14:08 wm-bot2: Test SAL log ([[phab:T325756|T325756]]) - cookbook ran by root@cloudcumin1001 * 14:07 wm-bot2: Test SAL log ([[phab:T325756|T325756]]) - cookbook ran by root@cloudcumin1001 === 2023-06-30 === * 21:37 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 21:34 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-06-29 === * 19:49 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:49 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:49 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:49 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:49 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:49 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:48 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:48 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:48 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:48 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:48 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:47 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:34 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:34 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:34 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:34 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:34 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['cinder-scheduler', 'cinder-volume'] - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['cinder-scheduler', 'cinder-volume'] - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:30 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:30 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-06-26 === * 08:46 arturo: [codfw1dev] manually start mariadb service @ cloudinfra-db-01.cloudinfra-codfw1dev.codfw1dev.wikimedia.cloud === 2023-06-24 === * 19:21 wm-bot2: Created new flavor: g3.cores16.ram34.disk20 (id:7dd33202-32c3-4bc7-b2d4-{{Gerrit|10c2ebe7e5c5}}) - cookbook ran by andrew@bullseye * 19:20 wm-bot2: Created new flavor: g3.cores16.ram34816.disk20.admin (id:d76925a8-b58b-489e-a3d9-{{Gerrit|c68c7744b551}}) - cookbook ran by andrew@bullseye === 2023-06-23 === * 16:45 wm-bot2: Restarting openstack services on cloudservices1005: ['designate-producer', 'designate-sink', 'designate-worker', 'designate-central', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 16:45 wm-bot2: Restarting openstack services on cloudservices1004: ['designate-worker', 'designate-api', 'designate-mdns', 'designate-producer', 'designate-central', 'designate-sink', 'designate-agent'] - cookbook ran by andrew@bullseye * 16:45 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] - cookbook ran by andrew@bullseye * 16:45 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] - cookbook ran by andrew@bullseye * 16:45 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:45 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:44 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:43 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:43 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:43 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:43 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:43 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:43 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:41 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 13:59 andrewbogott: rebooting every VM in codfw1dev === 2023-06-13 === * 15:35 wm-bot2: Restarting openstack services on cloudservices1005: ['designate-producer', 'designate-sink', 'designate-worker', 'designate-central', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudservices1004: ['designate-worker', 'designate-api', 'designate-mdns', 'designate-producer', 'designate-central', 'designate-sink', 'designate-agent'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:35 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:34 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:32 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:31 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:22 wm-bot2: Upgraded and rebooted host cloudcontrol1007.wikimedia.org - cookbook ran by andrew@bullseye * 15:11 wm-bot2: Upgraded and rebooted host cloudcontrol1006.wikimedia.org - cookbook ran by andrew@bullseye * 14:58 wm-bot2: Upgraded and rebooted host cloudcontrol1005.wikimedia.org - cookbook ran by andrew@bullseye === 2023-06-12 === * 19:37 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:37 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:37 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:37 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:37 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:37 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:37 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:35 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:35 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 11:57 arturo: [codfw1dev] refresh various occurrences of old FQDNs in instance puppet via horizon ([[phab:T324992|T324992]]) === 2023-06-08 === * 16:56 andrewbogott: deleting all Stretch base images from glance * 16:45 andrewbogott: updated the bullseye image with https://cloud.debian.org/images/cloud/bullseye/20230601-1398/debian-11-genericcloud-amd64-20230601-1398.tar.xz * 12:17 wm-bot2: Drained cloudvirt1047.eqiad.wmnet ([[phab:T334644|T334644]]) - cookbook ran by dcaro@vulcanus * 12:17 wm-bot2: Set cloudvirt cloudvirt1047.eqiad.wmnet maintenance (downtime id: 02920314-1efe-4934-ad81-{{Gerrit|d2a6cf2e17ab}}, use this to unset) ([[phab:T334644|T334644]]) - cookbook ran by dcaro@vulcanus * 12:16 wm-bot2: Draining cloudvirt1047.eqiad.wmnet ([[phab:T334644|T334644]]) - cookbook ran by dcaro@vulcanus * 12:06 wm-bot2: Set cloudvirt cloudvirt1047.eqiad.wmnet maintenance (downtime id: 769349bf-465f-4f0c-a8f3-{{Gerrit|f2423631ba7e}}, use this to unset) ([[phab:T334644|T334644]]) - cookbook ran by dcaro@vulcanus * 12:05 wm-bot2: Draining cloudvirt1047.eqiad.wmnet ([[phab:T334644|T334644]]) - cookbook ran by dcaro@vulcanus === 2023-06-07 === * 10:06 dcaro: upgraded ruby2.5 to latest fixed version on all buster VMs * 08:10 dcaro: downgrading ruby2.5 to previous backport on all buster VMs === 2023-06-06 === * 19:09 andrewbogott: also increased RAM and secgroup-rule quota for Trove [[phab:T337882|T337882]] * 19:06 andrewbogott: increased trove secgroups, instances, volumes quotas from 40 to 100. Trove is too popular! [[phab:T337882|T337882]] * 18:31 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by andrew@bullseye * 18:31 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:30 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:29 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:28 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:27 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:26 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute'] - cookbook ran by andrew@bullseye * 18:26 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:55 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:55 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:55 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:54 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:54 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:54 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:42 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 17:42 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 17:41 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 17:41 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['cinder-scheduler', 'cinder-volume'] - cookbook ran by andrew@bullseye * 17:41 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['cinder-scheduler', 'cinder-volume'] - cookbook ran by andrew@bullseye === 2023-06-05 === * 09:41 arturo: [codfw1dev] rebooting bastion-codfw1dev-02 (no IP address in the main interface) [[phab:T336963|T336963]] === 2023-06-02 === * 14:40 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by arturo@nostromo * 12:59 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 12:58 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 12:58 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 12:44 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 12:44 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 12:43 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-05-26 === * 16:03 andrewbogott: "maintain-views --all-databases --replace-all" on clouddb1021 for [[phab:T337446|T337446]] === 2023-05-22 === * 19:54 andrewbogott: deleting project 'citelearn' as per https://wikitech.wikimedia.org/wiki/News/Cloud_VPS_2022_Purge#SHUTDOWN_citelearn === 2023-05-21 === * 23:29 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 23:29 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 23:29 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 23:29 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 23:29 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 23:29 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 23:28 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 23:28 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 23:28 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 23:28 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 23:28 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 23:27 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:45 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:44 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 21:44 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-05-19 === * 14:53 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 14:53 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 14:53 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 14:53 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 14:53 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 14:53 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 14:53 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 14:52 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 14:52 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 14:52 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 14:52 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 14:52 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-05-18 === * 21:39 andrewbogott: deleting obsolete roles '4d8cad783d6342efa8414d7d36fbc034 {{!}} projectadmin_renamed_for_[[phab:T330759|T330759]]' and 'f473273fac7146b3bdbf22e5d4504f95 {{!}} user_renamed_for_[[phab:T330759|T330759]]' on eqiad1. State pre-deletion is dumped to /root/allassignmentspredeletion.txt on cloudcontrol1007. * 21:35 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 21:35 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 21:34 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 21:34 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:34 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:34 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 21:34 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 21:33 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:20 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:20 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 19:20 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:20 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:20 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:20 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 19:19 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:19 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:19 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:19 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 19:19 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 19:18 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:12 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 16:12 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 16:12 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:12 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:12 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:12 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:11 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:11 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:11 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:11 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:11 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:10 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:22 wm-bot2: Restarting openstack services on cloudcontrol2001-dev@local1: ['cinder-volume'] - cookbook ran by andrew@bullseye * 15:21 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 15:21 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 15:21 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:21 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:21 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:21 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:21 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:20 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-05-15 === * 14:28 wm-bot2: Drained cloudvirt1034.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:28 wm-bot2: Set cloudvirt cloudvirt1034.eqiad.wmnet maintenance (downtime id: 96ce2ed0-3aff-4d04-be0b-{{Gerrit|e16513070617}}, use this to unset) - cookbook ran by andrew@bullseye * 14:27 wm-bot2: Draining cloudvirt1034.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:23 wm-bot2: Drained cloudvirt1027.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 14:17 wm-bot2: Set cloudvirt cloudvirt1027.eqiad.wmnet maintenance (downtime id: 110176f8-04d5-4110-bb7d-{{Gerrit|1ab272bd8be2}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 14:16 wm-bot2: Draining cloudvirt1027.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 14:13 wm-bot2: Set cloudvirt cloudvirt1033.eqiad.wmnet maintenance (downtime id: c6e92e13-49f4-4db3-8a13-{{Gerrit|8692ccfd3bc9}}, use this to unset) - cookbook ran by andrew@bullseye * 14:12 wm-bot2: Draining cloudvirt1033.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:12 wm-bot2: Drained cloudvirt1035.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 14:11 wm-bot2: Set cloudvirt cloudvirt1033.eqiad.wmnet maintenance (downtime id: fa730dec-848f-45fb-9eda-{{Gerrit|e74bd874c5c9}}, use this to unset) - cookbook ran by andrew@bullseye * 14:10 wm-bot2: Draining cloudvirt1033.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:06 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by andrew@bullseye * 14:06 wm-bot2: Restarting openstack services on cloudcontrol1006: ['cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 14:06 wm-bot2: Restarting openstack services on cloudcontrol1007: ['cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 14:06 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by andrew@bullseye * 14:06 wm-bot2: Restarting openstack services on cloudcontrol1005: ['cinder-volume', 'cinder-scheduler'] - cookbook ran by andrew@bullseye * 14:01 wm-bot2: Set cloudvirt cloudvirt1033.eqiad.wmnet maintenance (downtime id: eb1cfac0-d481-4baa-b9cd-{{Gerrit|15e5fbcef495}}, use this to unset) - cookbook ran by andrew@bullseye * 14:00 wm-bot2: Draining cloudvirt1033.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:58 wm-bot2: Set cloudvirt cloudvirt1034.eqiad.wmnet maintenance (downtime id: 3e6c3ff3-7d55-4777-9032-{{Gerrit|b867a257eced}}, use this to unset) - cookbook ran by andrew@bullseye * 13:57 wm-bot2: Draining cloudvirt1034.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:53 wm-bot2: Set cloudvirt cloudvirt1033.eqiad.wmnet maintenance (downtime id: 0693664a-df78-417e-ba34-{{Gerrit|590e5a0a9981}}, use this to unset) - cookbook ran by andrew@bullseye * 13:52 wm-bot2: Draining cloudvirt1033.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:49 wm-bot2: Set cloudvirt cloudvirt1035.eqiad.wmnet maintenance (downtime id: e6929ab8-4bc3-4186-817b-{{Gerrit|9b53dbd597c6}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:48 wm-bot2: Draining cloudvirt1035.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:40 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:40 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:40 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:40 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:40 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:40 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:39 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:38 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:37 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:36 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 13:36 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:36 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:36 andrewbogott: restarting nova services in eqiad1, trying to free up db connections * 13:36 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:36 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:36 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:36 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute'] - cookbook ran by andrew@bullseye * 13:33 wm-bot2: Set cloudvirt cloudvirt1034.eqiad.wmnet maintenance (downtime id: fef43d73-6fd4-4dde-a0ac-{{Gerrit|95fd69a9b0c1}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:32 wm-bot2: Draining cloudvirt1034.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:29 wm-bot2: Draining cloudvirt1027.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:24 wm-bot2: Set cloudvirt cloudvirt1027.eqiad.wmnet maintenance (downtime id: 5f867662-e824-498c-a715-{{Gerrit|2e2ad50f0bb5}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:23 wm-bot2: Draining cloudvirt1027.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:08 wm-bot2: Set cloudvirt cloudvirt1033.eqiad.wmnet maintenance (downtime id: 0f515150-1313-41d4-a5f6-{{Gerrit|9bc00ce9b245}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:07 wm-bot2: Draining cloudvirt1033.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 13:07 wm-bot2: Drained cloudvirt1032.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:50 wm-bot2: Set cloudvirt cloudvirt1032.eqiad.wmnet maintenance (downtime id: e826adc3-addd-44d8-b39e-{{Gerrit|ae7bd2df1e60}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:49 wm-bot2: Draining cloudvirt1032.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:49 wm-bot2: Drained cloudvirt1031.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:14 wm-bot2: Set cloudvirt cloudvirt1027.eqiad.wmnet maintenance (downtime id: 4154c818-744c-4d84-9883-{{Gerrit|cae7a5826ed5}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:13 wm-bot2: Draining cloudvirt1027.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:13 wm-bot2: Drained cloudvirt1026.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:04 wm-bot2: Set cloudvirt cloudvirt1026.eqiad.wmnet maintenance (downtime id: cdfc3d01-ec1e-483c-9a38-{{Gerrit|834193e487ff}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:03 wm-bot2: Draining cloudvirt1026.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 12:03 wm-bot2: Drained cloudvirt1025.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 11:51 wm-bot2: Set cloudvirt cloudvirt1025.eqiad.wmnet maintenance (downtime id: 6a56757d-35de-499e-8209-{{Gerrit|728bcf62a22a}}, use this to unset) ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus * 11:50 wm-bot2: Draining cloudvirt1025.eqiad.wmnet ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus === 2023-05-12 === * 17:52 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-05-10 === * 16:09 wm-bot2: Restarting openstack services on cloudservices1005: ['designate-producer', 'designate-sink', 'designate-worker', 'designate-central', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 16:09 wm-bot2: Restarting openstack services on cloudservices1004: ['designate-worker', 'designate-api', 'designate-mdns', 'designate-producer', 'designate-central', 'designate-sink', 'designate-agent'] - cookbook ran by andrew@bullseye * 16:09 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] - cookbook ran by andrew@bullseye * 16:09 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] - cookbook ran by andrew@bullseye * 16:09 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:09 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:08 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:07 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:07 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:07 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:07 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:07 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:07 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:06 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:05 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:04 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:04 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:04 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1024: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 16:00 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:59 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:58 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:57 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:57 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:57 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:57 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:57 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:57 wm-bot2: Restarting openstack services on cloudvirt1024: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:57 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudbackup2001: ['cinder-backup'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudbackup2002: ['cinder-backup'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:56 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:55 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:55 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:55 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:55 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:55 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:55 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:55 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:54 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:53 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:52 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 15:52 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:52 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:52 andrewbogott: running "cookbook -c ~/.config/spicerack/cookbook_config.yaml wmcs.openstack.restart_openstack --cluster-name eqiad1 --all" to pick up changes for testing [[phab:T336379|T336379]] * 15:52 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:52 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:52 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 15:52 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 03:05 andrewbogott: "systemctl restart puppet-enc" on enc-1.cloudinfra.eqiad1.wikimedia.cloud. Seems to have crashed. === 2023-05-05 === * 16:07 wm-bot2: Drained cloudvirt1024.eqiad.wmnet ([[phab:T336064|T336064]]) - cookbook ran by andrew@bullseye * 16:03 wm-bot2: Set cloudvirt cloudvirt1024.eqiad.wmnet maintenance (downtime id: 95995009-09d6-496e-8cd2-{{Gerrit|0cfac93d3cf7}}, use this to unset) ([[phab:T336064|T336064]]) - cookbook ran by andrew@bullseye * 16:02 wm-bot2: Draining cloudvirt1024.eqiad.wmnet ([[phab:T336064|T336064]]) - cookbook ran by andrew@bullseye * 16:01 wm-bot2: Drained cloudvirt1023.eqiad.wmnet ([[phab:T336064|T336064]]) - cookbook ran by andrew@bullseye * 15:51 wm-bot2: Set cloudvirt cloudvirt1023.eqiad.wmnet maintenance (downtime id: 53c46cae-00af-4664-97ff-{{Gerrit|266b393335bb}}, use this to unset) ([[phab:T336064|T336064]]) - cookbook ran by andrew@bullseye * 15:50 wm-bot2: Draining cloudvirt1023.eqiad.wmnet ([[phab:T336064|T336064]]) - cookbook ran by andrew@bullseye * 15:49 wm-bot2: Set cloudvirt cloudvirt1024.eqiad.wmnet maintenance (downtime id: 528ea4f6-8088-475e-937f-{{Gerrit|098ffba861b6}}, use this to unset) - cookbook ran by andrew@bullseye * 15:47 wm-bot2: Set cloudvirt cloudvirt1023.eqiad.wmnet maintenance (downtime id: 3ef85b5e-d9d9-4b24-901b-{{Gerrit|a3058a7d0615}}, use this to unset) - cookbook ran by andrew@bullseye * 15:44 andrewbogott: moved cloudvirt1023 and cloudvirt1024 from 'ceph' aggregate to 'maintenance' aggregate, prep for decom [[phab:T336064|T336064]] * 15:44 andrewbogott: moved cloudvirt1028 from 'localdisk' aggregate to 'maintenance' aggregate. Nothing new should be scheduled here, local storage should now move to cloudvirtlocal100x * 15:41 andrewbogott: moved cloudvirt1055 and cloudvirt1056 from 'spare' to 'ceph' aggregate. Prep for removing two obsolete cloudvirts, 1023 and 1024. [[phab:T336064|T336064]] === 2023-05-04 === * 22:49 andrewbogott: removed fullstack-* puppet reports on puppetmaster-02.cloudinfra-codfw1dev.codfw1dev.wikimedia.cloud and cloud-puppetmaster-03.cloudinfra.eqiad.wmflabs to free up disk space === 2023-05-02 === * 13:01 wm-bot2: Adding OSD cloudcephosd2001-dev.codfw.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 13:01 wm-bot2: Adding new OSDs ['cloudcephosd2001-dev.codfw.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 12:31 wm-bot2: Destroying OSDs with ids in [0] on cloudcephosd2001-dev from codfw1 - cookbook ran by dcaro@vulcanus * 12:30 wm-bot2: Depooling OSDs with ids in [0] on cloudcephosd2001-dev from codfw1 - cookbook ran by dcaro@vulcanus * 11:53 wm-bot2: The cluster is now rebalanced after adding the new OSDs ['cloudcephosd2001-dev.codfw.wmnet'] - cookbook ran by dcaro@vulcanus * 11:53 wm-bot2: Added 1 new OSDs ['cloudcephosd2001-dev.codfw.wmnet'] - cookbook ran by dcaro@vulcanus * 11:53 wm-bot2: Added OSD cloudcephosd2001-dev.codfw.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 11:53 wm-bot2: Adding OSD cloudcephosd2001-dev.codfw.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 11:52 wm-bot2: Adding new OSDs ['cloudcephosd2001-dev.codfw.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus === 2023-05-01 === * 17:09 wm-bot2: Adding OSD cloudcephosd2001-dev.codfw.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 17:09 wm-bot2: Adding new OSDs ['cloudcephosd2001-dev.codfw.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 17:08 wm-bot2: Adding OSD cloudcephosd2001-dev.codfw.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 17:08 wm-bot2: Adding new OSDs ['cloudcephosd2001-dev.codfw.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 15:22 wm-bot2: Depooling OSDs with ids in [0] on cloudcephosd2001-dev from codfw1 - cookbook ran by dcaro@vulcanus * 13:53 taavi: running wmcs-novastats-puppetleaks in real mode [[phab:T334127|T334127]] * 13:51 wm-bot2: Depooling OSDs with ids in [0] on cloudcephosd2001-dev from codfw1 - cookbook ran by dcaro@vulcanus === 2023-04-18 === * 22:52 wm-bot2: Restarting openstack services on cloudcontrol1007: ['neutron-api', 'neutron-rpc-server'] - cookbook ran by andrew@bullseye * 22:52 wm-bot2: Restarting openstack services on cloudcontrol1006: ['neutron-api', 'neutron-rpc-server'] - cookbook ran by andrew@bullseye * 22:52 wm-bot2: Restarting openstack services on cloudcontrol1005: ['neutron-api', 'neutron-rpc-server'] - cookbook ran by andrew@bullseye * 22:52 wm-bot2: Restarting openstack services on cloudvirt1056: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1019: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1060: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1023: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1038: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1036: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1039: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1050: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1061: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1028: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1020: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1052: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1040: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1043: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1026: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1030: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1048: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1025: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Restarting openstack services on cloudvirt1035: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1055: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1042: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1059: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1057: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1032: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1024: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1029: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1044: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1047: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1034: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1058: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1045: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1041: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1031: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1046: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1027: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudnet1006: ['neutron-linuxbridge-agent', 'neutron-metadata-agent', 'neutron-dhcp-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudnet1005: ['neutron-linuxbridge-agent', 'neutron-dhcp-agent', 'neutron-metadata-agent'] - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Restarting openstack services on cloudvirt1054: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:49 wm-bot2: Restarting openstack services on cloudvirt1033: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:49 wm-bot2: Restarting openstack services on cloudvirt1037: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:49 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:49 wm-bot2: Restarting openstack services on cloudvirt1051: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:49 wm-bot2: Restarting openstack services on cloudvirt1053: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:49 wm-bot2: Restarting openstack services on cloudvirt1049: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 22:48 wm-bot2: Restarting openstack services on cloudvirtlocal1003: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:48 wm-bot2: Restarting openstack services on cloudvirtlocal1002: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:48 wm-bot2: Restarting openstack services on cloudvirtlocal1001: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:48 wm-bot2: Restarting openstack services on cloudvirt1054: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:48 wm-bot2: Restarting openstack services on cloudvirt1055: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:48 wm-bot2: Restarting openstack services on cloudvirt1060: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1058: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1059: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1061: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1057: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1056: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1051: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1050: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1049: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 22:47 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1019: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1020: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:45 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:44 wm-bot2: Restarting openstack services on cloudvirt1023: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:44 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:44 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:44 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:44 wm-bot2: Restarting openstack services on cloudvirt1024: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:41 andrewbogott: resetting rabbitmq on cloudrabbit1003 due to splitbrain * 22:40 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:40 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:40 wm-bot2: Restarting openstack services on cloudvirt1023: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:39 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:39 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:39 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:39 wm-bot2: Restarting openstack services on cloudvirt1024: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:36 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:36 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:36 wm-bot2: Restarting openstack services on cloudvirt1024: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:34 wm-bot2: Restarting openstack services on cloudvirt1053: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:34 wm-bot2: Restarting openstack services on cloudvirt1052: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:34 wm-bot2: Restarting openstack services on cloudvirt1048: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Restarting openstack services on cloudcontrol1007: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Restarting openstack services on cloudcontrol1006: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Restarting openstack services on cloudvirt-wdqs1003: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Restarting openstack services on cloudvirt-wdqs1002: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Restarting openstack services on cloudvirt-wdqs1001: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Restarting openstack services on cloudvirt1047: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Restarting openstack services on cloudvirt1038: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:32 wm-bot2: Restarting openstack services on cloudvirt1042: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:32 wm-bot2: Restarting openstack services on cloudvirt1044: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:32 wm-bot2: Restarting openstack services on cloudvirt1041: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:31 wm-bot2: Restarting openstack services on cloudvirt1046: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:31 wm-bot2: Restarting openstack services on cloudvirt1043: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:31 wm-bot2: Restarting openstack services on cloudvirt1045: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:31 wm-bot2: Restarting openstack services on cloudvirt1040: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Restarting openstack services on cloudvirt1036: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Restarting openstack services on cloudvirt1034: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Restarting openstack services on cloudvirt1039: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Restarting openstack services on cloudvirt1037: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Restarting openstack services on cloudvirt1035: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Restarting openstack services on cloudvirt1033: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1031: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1032: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudcontrol1005: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1028: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1019: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1020: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1030: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1027: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:29 wm-bot2: Restarting openstack services on cloudvirt1023: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:28 wm-bot2: Restarting openstack services on cloudvirt1026: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:28 wm-bot2: Restarting openstack services on cloudvirt1029: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:28 wm-bot2: Restarting openstack services on cloudvirt1025: ['nova-compute'] - cookbook ran by andrew@bullseye * 22:28 wm-bot2: Restarting openstack services on cloudvirt1024: ['nova-compute'] - cookbook ran by andrew@bullseye === 2023-04-17 === * 08:49 wm-bot2: Increased quotas by 9 cores, 1 instances, 16 ram ([[phab:T334695|T334695]]) - cookbook ran by dcaro@vulcanus === 2023-04-06 === * 17:03 andrewbogott: running wmcs-wikireplica-dns on cloudcontrol1005 to update tools-db dns entries === 2023-04-04 === * 17:23 andrewbogott: resetting all three rabbitmq nodes and restarting all openstack services as per https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Rabbitmq#Resetting_the_HA_setup === 2023-03-28 === * 13:56 andrewbogott: depooling cloudweb1003 before switch upgrade * 10:58 dhinus: disabled tool "wb" by clicking the disable button at https://toolsadmin.wikimedia.org/tools/id/wb [[phab:T328693|T328693]] * 08:34 arturo: cleanup neutron agents for cloudvirt1021/1022 (decom) * 08:32 arturo: cleanup neutron agents for cloudvirt1017 (decom) === 2023-03-27 === * 14:30 wm-bot2: Drained cloudvirt1024.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:19 wm-bot2: Set cloudvirt cloudvirt1024.eqiad.wmnet maintenance (downtime id: 3f43d3ca-696c-4d3c-8d5b-{{Gerrit|e57984f0eb86}}, use this to unset) - cookbook ran by andrew@bullseye * 14:18 wm-bot2: Draining cloudvirt1024.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:17 wm-bot2: Drained cloudvirt1023.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:08 wm-bot2: Set cloudvirt cloudvirt1023.eqiad.wmnet maintenance (downtime id: f4767781-ea26-453b-9521-{{Gerrit|847a8340a249}}, use this to unset) - cookbook ran by andrew@bullseye * 14:07 wm-bot2: Draining cloudvirt1023.eqiad.wmnet - cookbook ran by andrew@bullseye * 14:04 wm-bot2: Drained cloudvirt1022.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:55 wm-bot2: Set cloudvirt cloudvirt1022.eqiad.wmnet maintenance (downtime id: 0e794cfe-5896-46c1-842d-{{Gerrit|c34719140d4f}}, use this to unset) - cookbook ran by andrew@bullseye * 13:54 wm-bot2: Draining cloudvirt1022.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:54 wm-bot2: Drained cloudvirt1021.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:48 wm-bot2: Set cloudvirt cloudvirt1021.eqiad.wmnet maintenance (downtime id: e7a904ca-003d-450a-ad85-{{Gerrit|886fb80dfc41}}, use this to unset) - cookbook ran by andrew@bullseye * 13:47 wm-bot2: Draining cloudvirt1021.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:46 wm-bot2: Drained cloudvirt1017.eqiad.wmnet - cookbook ran by andrew@bullseye * 13:36 wm-bot2: Set cloudvirt cloudvirt1017.eqiad.wmnet maintenance (downtime id: 0ff0090f-26c3-4278-a9c9-{{Gerrit|cce518559408}}, use this to unset) - cookbook ran by andrew@bullseye * 13:35 wm-bot2: Draining cloudvirt1017.eqiad.wmnet - cookbook ran by andrew@bullseye === 2023-03-22 === * 12:41 taavi: delete wmde-templates-alpha project [[phab:T332773|T332773]] === 2023-03-08 === * 21:45 bd808: maintain-kubeusers container in CrashLoopBackoff, investigating * 13:49 dcaro: stopping puppet on labostre1004 to debug maintain-dbusers === 2023-03-07 === * 16:06 andrewbogott: updated application credential roles, replacing 'user' with 'reader' and 'projectadmin' with 'member': update application_credential_role set role_id='f75a3c410bca4e96a1cf6ac103b0ccaf' where role_id='f473273fac7146b3bdbf22e5d4504f95' and update application_credential_role set role_id='38676f30eaeb44518bf7e144a73c8da6' where role_id='4d8cad783d6342efa8414d7d36fbc034' * 10:11 dcaro: there was a little unavailability for some VMs while ceph was starting to rebalance things, but it seems stable and moving data around ([[phab:T331141|T331141]]) * 09:37 dcaro: Changing ceph crush map to allow rack HA on eqiad1 cluster ([[phab:T331141|T331141]]) === 2023-03-03 === * 12:12 arturo: installing haproxy updates ([[phab:T331119|T331119]]) === 2023-03-02 === * 16:17 wm-bot2: The cluster is now rebalanced after adding the new OSDs ['cloudcephosd1010.eqiad.wmnet'] ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 14:21 wm-bot2: Added 1 new OSDs ['cloudcephosd1010.eqiad.wmnet'] ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 14:21 wm-bot2: Added OSD cloudcephosd1010.eqiad.wmnet... (1/1) ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 14:13 wm-bot2: Finished rebooting node cloudcephosd1010.eqiad.wmnet ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 14:10 wm-bot2: Rebooting node cloudcephosd1010.eqiad.wmnet ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 14:09 wm-bot2: Adding OSD cloudcephosd1010.eqiad.wmnet... (1/1) ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 14:09 wm-bot2: Adding new OSDs ['cloudcephosd1010.eqiad.wmnet'] to the cluster ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 10:43 wm-bot2: The cluster is now rebalanced after adding the new OSDs ['cloudcephosd1005.eqiad.wmnet'] ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 08:44 wm-bot2: Added 1 new OSDs ['cloudcephosd1005.eqiad.wmnet'] ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 08:44 wm-bot2: Added OSD cloudcephosd1005.eqiad.wmnet... (1/1) ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 08:36 wm-bot2: Finished rebooting node cloudcephosd1005.eqiad.wmnet ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 08:33 wm-bot2: Rebooting node cloudcephosd1005.eqiad.wmnet ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 08:32 wm-bot2: Adding OSD cloudcephosd1005.eqiad.wmnet... (1/1) ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 08:32 wm-bot2: Adding new OSDs ['cloudcephosd1005.eqiad.wmnet'] to the cluster ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus === 2023-02-28 === * 22:47 andrewbogott: adding new 'member' role assignment to every user/project pair that currently has the 'user' assignment. [[phab:T330759|T330759]] * 09:47 wm-bot2: Depooled and destroyed OSD daemons [79, 78, 77, 76, 75, 74, 73, 72] and removed the OSD host cloudcephosd1010 from the CRUSH map. ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 09:46 wm-bot2: Destroying OSDs with ids in [79, 78, 77, 76, 75, 74, 73, 72] on cloudcephosd1010 from eqiad1 ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 09:28 wm-bot2: Depooling OSDs with ids in [79, 78, 77, 76, 75, 74, 73, 72] on cloudcephosd1010 from eqiad1 ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 09:13 wm-bot2: Depooling OSDs with ids in [79, 78, 77, 76, 75, 74, 73, 72] on cloudcephosd1010 from eqiad1 ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 09:12 wm-bot2: Depooled and destroyed OSD daemons [39, 38, 37, 36, 35, 34, 33, 32] and removed the OSD host cloudcephosd1005 from the CRUSH map. ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 09:11 wm-bot2: Destroying OSDs with ids in [39, 38, 37, 36, 35, 34, 33, 32] on cloudcephosd1005 from eqiad1 ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus * 08:55 wm-bot2: Depooling OSDs with ids in [39, 38, 37, 36, 35, 34, 33, 32] on cloudcephosd1005 from eqiad1 ([[phab:T329504|T329504]]) - cookbook ran by dcaro@vulcanus === 2023-02-27 === * 21:01 wm-bot2: The cluster is now rebalanced after adding the new OSDs ['cloudcephosd1004.eqiad.wmnet'] ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:53 wm-bot2: Added 1 new OSDs ['cloudcephosd1004.eqiad.wmnet'] ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:53 wm-bot2: Added OSD cloudcephosd1004.eqiad.wmnet... (1/1) ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:45 wm-bot2: Finished rebooting node cloudcephosd1004.eqiad.wmnet ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:42 wm-bot2: Rebooting node cloudcephosd1004.eqiad.wmnet ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:41 wm-bot2: Adding OSD cloudcephosd1004.eqiad.wmnet... (1/1) ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:41 wm-bot2: Adding new OSDs ['cloudcephosd1004.eqiad.wmnet'] to the cluster ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:38 wm-bot2: Rebooting node cloudcephosd1004.eqiad.wmnet ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:38 wm-bot2: Adding OSD cloudcephosd1004.eqiad.wmnet... (1/1) ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:38 wm-bot2: Adding new OSDs ['cloudcephosd1004.eqiad.wmnet'] to the cluster ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 18:14 wm-bot2: The cluster is now rebalanced after adding the new OSDs ['cloudcephosd1003.eqiad.wmnet'] ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 16:13 wm-bot2: Added 1 new OSDs ['cloudcephosd1003.eqiad.wmnet'] ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 16:13 wm-bot2: Added OSD cloudcephosd1003.eqiad.wmnet... (1/1) ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 16:05 wm-bot2: Finished rebooting node cloudcephosd1003.eqiad.wmnet ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 16:02 wm-bot2: Rebooting node cloudcephosd1003.eqiad.wmnet ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 16:01 wm-bot2: Adding OSD cloudcephosd1003.eqiad.wmnet... (1/1) ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 16:01 wm-bot2: Adding new OSDs ['cloudcephosd1003.eqiad.wmnet'] to the cluster ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 15:59 wm-bot2: Adding OSD coludcephosd1003.eqiad.wmnet... (1/1) ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 15:59 wm-bot2: Adding new OSDs ['coludcephosd1003.eqiad.wmnet'] to the cluster ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 15:58 wm-bot2: Adding OSD coludcephosd1003... (1/1) ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 15:58 wm-bot2: Adding new OSDs ['coludcephosd1003'] to the cluster ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus === 2023-02-22 === * 08:25 wm-bot2: Depooled and destroyed OSD daemons [31, 30, 29, 28, 27, 26, 25, 24] and removed the OSD host cloudcephosd1004 from the CRUSH map. ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 08:25 wm-bot2: Destroying OSDs with ids in [31, 30, 29, 28, 27, 26, 25, 24] on cloudcephosd1004 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 08:22 wm-bot2: Depooling OSDs with ids in [31, 30, 29, 28, 27, 26, 25, 24] on cloudcephosd1004 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 08:15 wm-bot2: Destroying OSDs with ids in [31, 30, 29, 28, 27, 26, 25, 24] on cloudcephosd1004 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 08:13 wm-bot2: Depooling OSDs with ids in [31, 30, 29, 28, 27, 26, 25, 24] on cloudcephosd1004 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 07:23 wm-bot2: Destroying OSDs with ids in [31, 30, 29, 28, 27, 26, 25, 24] on cloudcephosd1004 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 07:02 wm-bot2: Depooling OSDs with ids in [31, 30, 29, 28, 27, 26, 25, 24] on cloudcephosd1004 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus === 2023-02-21 === * 20:27 andrewbogott: deleted 200 more orphaned VM images with wmcs-novastats-cephleaks * 17:18 andrewbogott: shutting down postgres on clouddb1004/1003, then shutting down the vms * 15:50 wm-bot2: Destroying OSDs with ids in [71, 70, 69, 68, 67, 66, 65, 64] on cloudcephosd1003 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 15:48 wm-bot2: Depooling OSDs with ids in [71, 70, 69, 68, 67, 66, 65, 64] on cloudcephosd1003 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 14:21 wm-bot2: Destroying OSDs with ids in [71, 70, 69, 68, 67, 66, 65, 64] on cloudcephosd1003 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus * 14:00 wm-bot2: Depooling OSDs with ids in [71, 70, 69, 68, 67, 66, 65, 64] on cloudcephosd1003 from eqiad1 ([[phab:T329502|T329502]]) - cookbook ran by dcaro@vulcanus === 2023-02-16 === * 19:14 wm-bot2: The cluster is now rebalanced after adding the new OSDs ['cloudcephosd1002.eqiad.wmnet'] ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 17:55 dcaro: Manually zapped /dev/sdc on cloudcephosd1002, probably a leftover drive since the beginning (or during the reimage the drives changed names, and this one had leftovers from the previous OS) ([[phab:T329498|T329498]]) * 17:47 wm-bot2: Added 1 new OSDs ['cloudcephosd1002.eqiad.wmnet'] ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 17:47 wm-bot2: Added OSD cloudcephosd1002.eqiad.wmnet... (1/1) ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 17:42 wm-bot2: Adding OSD cloudcephosd1002.eqiad.wmnet... (1/1) ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 17:42 wm-bot2: Adding new OSDs ['cloudcephosd1002.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 16:50 wm-bot2: Adding OSD cloudcephosd1002.eqiad.wmnet... (1/1) ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 16:50 wm-bot2: Adding new OSDs ['cloudcephosd1002.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 16:01 wm-bot2: Adding OSD cloudcephosd1002.eqiad.wmnet... (1/1) ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 16:01 wm-bot2: Adding new OSDs ['cloudcephosd1002.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 16:00 wm-bot2: Adding OSD cloudcephosd1002.eqiad.wmnet... (1/1) ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 16:00 wm-bot2: Adding new OSDs ['cloudcephosd1002.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 14:05 wm-bot2: Added 1 new OSDs ['cloudcephosd1001.eqiad.wmnet'] ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:29 wm-bot2: Adding new OSDs ['cloudcephosd1001.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:29 wm-bot2: Adding new OSDs ['cloudcephosd1001.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:24 wm-bot2: Adding OSD cloudcephosd1001.eqiad.wmnet... (1/1) ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:24 wm-bot2: Adding new OSDs ['cloudcephosd1001.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:23 wm-bot2: Destroying OSDs with ids in [63, 62, 61, 60, 59, 58, 57, 56] on cloudcephosd1002 from eqiad1 ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:21 wm-bot2: Depooling OSDs with ids in [63, 62, 61, 60, 59, 58, 57, 56] on cloudcephosd1002 from eqiad1 ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:14 wm-bot2: Adding OSD cloudcephosd1001.eqiad.wmnet... (1/1) ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:14 wm-bot2: Adding new OSDs ['cloudcephosd1001.eqiad.wmnet'] to the cluster ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 11:20 wm-bot2: Destroying OSDs with ids in [53, 52, 51, 50] on cloudcephosd1001 from eqiad1 ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 11:19 wm-bot2: Depooling OSDs with ids in [53, 52, 51, 50] on cloudcephosd1001 from eqiad1 ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 11:03 wm-bot2: Destroying OSDs with ids in [55, 54, 53, 52, 51, 50] on cloudcephosd1001 from eqiad1 ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 11:01 wm-bot2: Depooling OSDs with ids in [55, 54, 53, 52, 51, 50] on cloudcephosd1001 from eqiad1 ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 10:59 wm-bot2: Depooling OSDs with ids in [55, 54, 53, 52, 51, 50] on cloudcephosd1001 from eqiad1 ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 10:15 dcaro: purges osd daemons 48 and 40 from eqiad ceph cluster ([[phab:T329709|T329709]]) === 2023-02-15 === * 14:53 andrewbogott: deleting another 100 leaked VM images with wmcs-novastats-cephleaks * 13:39 wm-bot2: Destroying OSDs with id [48] on cloudcephosd1001 from eqiad1 - cookbook ran by dcaro@vulcanus * 13:14 wm-bot2: Destroying OSDs with id [48] on cloudcephosd1001 from eqiad1 - cookbook ran by dcaro@vulcanus * 13:13 wm-bot2: Destroying OSDs with id [12345] on cloudcephosd1001 from eqiad1 - cookbook ran by dcaro@vulcanus * 13:11 wm-bot2: Destroying OSDs with id [12345] on cloudcephosd1001 from eqiad1 - cookbook ran by dcaro@vulcanus * 13:11 wm-bot2: Destroying OSDs with id [12345] on cloudcephosd1001 from eqiad1 - cookbook ran by dcaro@vulcanus * 13:10 wm-bot2: Destroying OSDs with id [[12345]] on cloudcephosd1001 from eqiad1 - cookbook ran by dcaro@vulcanus * 13:09 wm-bot2: Destroying OSDs with id ['12345'] on cloudcephosd1001 from eqiad1 - cookbook ran by dcaro@vulcanus === 2023-02-14 === * 13:17 andrewbogott: restarting all eqiad1 openstack services because that seems to sometimes help things *shrug* === 2023-02-13 === * 14:06 wm-bot2: Set the ceph cluster for eqiad1 in maintenance, alert silence ids: 8fbf6bfd-eec1-4d81-8e0d-{{Gerrit|ea431d8411ee}} ([[phab:T329498|T329498]]) - cookbook ran by dcaro@vulcanus * 13:32 taavi: re-enable puppet on labstore1004 [[phab:T329377|T329377]] === 2023-02-09 === * 21:17 andrewbogott: deleted 10% of leaked VM ceph images using wmcs-novastats-cephleaks (only 10% out of an abundance of caution) === 2023-02-08 === * 17:08 arturo: changing to cloudgw network setup, make VIPs /32 ([[phab:T295774|T295774]]) === 2023-02-07 === * 11:26 arturo: [codfw1dev] testing network changes in cloudgw, expect unrealiable network ([[phab:T295774|T295774]]) === 2023-02-04 === * 13:44 taavi: drop old columns from oathauth_users table on labtestwiki [[phab:T328131|T328131]] === 2023-02-03 === * 15:00 andrewbogott: restarted nova services in eqiad1 in an attempt to eke out another day or two of stability * 14:13 taavi: attached GrapheSuppression developer account to wikitech === 2023-02-02 === * 13:14 dcaro_away: draining osd.48 from node cloudcephosd1001 ([[phab:T316544|T316544]]) * 12:57 wm-bot2: Set the ceph cluster for eqiad1 in maintenance, alert silence ids: 7ac2b25a-d1bb-4789-8aa6-{{Gerrit|b9435b505349}} ([[phab:T316544|T316544]]) - cookbook ran by dcaro@vulcanus === 2023-01-30 === * 22:34 wm-bot2: Upgraded and rebooted host cloudrabbit1002.wikimedia.org - cookbook ran by andrew@bullseye * 21:34 andrewbogott: merging https://gerrit.wikimedia.org/r/c/operations/puppet/+/884922 and upgrading rabbitmq nodes for [[phab:T328155|T328155]] === 2023-01-27 === * 20:08 wm-bot2: Upgraded and rebooted host cloudcontrol2005-dev.wikimedia.org - cookbook ran by andrew@bullseye * 19:22 wm-bot2: Upgraded and rebooted host cloudcontrol2004-dev.wikimedia.org - cookbook ran by andrew@bullseye * 19:10 wm-bot2: Upgraded and rebooted host cloudcontrol2001-dev.wikimedia.org - cookbook ran by andrew@bullseye * 15:25 andrewbogott: restarting openstack services in eqiad1, another attempt to address instability === 2023-01-26 === * 20:34 andrewbogott: shutting down mariadb on cloudbackup2001-dev, testing the waters for [[phab:T328079|T328079]] === 2023-01-22 === * 03:42 andrewbogott: reset eqiad1 rabbitmq in an attempt to resolve some mild instability === 2023-01-20 === * 15:26 wm-bot2: Removed cloudweb hosts (cloudweb2002-dev.wikimedia.org) from maintenance mode. - cookbook ran by andrew@bullseye * 15:26 wm-bot2: Put cloudweb hosts (cloudweb2002-dev.wikimedia.org) into maintenance mode (downtime id: ['f47a3d91-b270-4c90-acc8-d85075a6bf8e'], use this to unset) - cookbook ran by andrew@bullseye * 13:15 arturo: reinstall python3-neutron (to reset manual patching) on all cloudnet nodes and patch it via puppet, then restart neutron-l3-agent by hand ([[phab:T327463|T327463]]) * 10:12 arturo: [codfw1dev] failover neutron-l3-agent between cloudnet2005-dev/cloudnet2006-dev a couple of times [[phab:T327463|T327463]] * 02:17 andrewbogott: stopping neutron-l3-agent on cloudnet1005 because it's logging at a furious rate and about to fill the drive === 2023-01-19 === * 18:06 wm-bot2: Removed cloudweb hosts (cloudweb2002-dev.wikimedia.org) from maintenance mode. - cookbook ran by andrew@bullseye * 18:06 wm-bot2: Put cloudweb hosts (cloudweb2002-dev.wikimedia.org) into maintenance mode (downtime id: ['36d5af6a-7d8e-4d0c-831e-1bf05c255984'], use this to unset) - cookbook ran by andrew@bullseye * 17:35 wm-bot2: Removed cloudweb hosts (cloudweb2002-dev.wikimedia.org) from maintenance mode. - cookbook ran by andrew@bullseye * 17:22 wm-bot2: Put cloudweb hosts (cloudweb2002-dev.wikimedia.org) into maintenance mode (downtime id: ['66ec4f04-d25b-4067-be9e-2fe12cb1d3ff'], use this to unset) - cookbook ran by andrew@bullseye === 2023-01-18 === * 22:32 wm-bot2: Set cloudweb cloudweb2002-dev.wikimedia.org maintenance (downtime id: 347cb75e-215e-4b85-ae14-{{Gerrit|4ce1934c70c7}}, use this to unset) - cookbook ran by andrew@bullseye * 22:01 wm-bot2: Set cloudweb cloudweb2002-dev.wikimedia.org maintenance (downtime id: 9bf6212f-7fdb-4869-8190-{{Gerrit|b07387b2bc7e}}, use this to unset) - cookbook ran by andrew@bullseye * 21:40 wm-bot2: Set cloudweb cloudweb2002-dev.wikimedia.org maintenance (downtime id: 11aec8ea-f443-41ae-b79a-{{Gerrit|5e1e3aa94546}}, use this to unset) - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 20:19 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 20:19 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 20:19 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 20:19 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 20:19 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 20:18 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 20:18 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 20:18 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 17:03 wm-bot2: Restarting openstack services on cloudservices2004-dev: ['designate-mdns', 'designate-sink', 'designate-central', 'designate-producer', 'designate-worker', 'designate-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudservices2005-dev: ['designate-central', 'designate-sink', 'designate-worker', 'designate-producer', 'designate-mdns', 'designate-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudbackup1002-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudbackup1001-dev: ['cinder-backup'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-volume', 'cinder-scheduler', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 17:02 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 17:01 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 14:38 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 14:38 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 14:37 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye * 14:37 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata', 'cinder-scheduler', 'cinder-volume', 'neutron-api', 'neutron-rpc-server', 'trove-api', 'trove-conductor', 'trove-taskmanager', 'keystone', 'keystone-admin', 'glance-api', 'magnum-api', 'magnum-conductor', 'heat-api', 'heat-api-cfn', 'heat-engine'] - cookbook ran by andrew@bullseye === 2023-01-17 === * 19:32 wm-bot2: Upgraded and rebooted host cloudbackup2002.codfw.wmnet - cookbook ran by andrew@bullseye * 18:04 wm-bot2: Upgraded and rebooted host cloudnet1005.eqiad.wmnet - cookbook ran by andrew@bullseye * 17:52 wm-bot2: Upgraded and rebooted host cloudcontrol1007.wikimedia.org - cookbook ran by andrew@bullseye * 17:36 wm-bot2: Upgraded and rebooted host cloudcontrol1006.wikimedia.org - cookbook ran by andrew@bullseye * 17:23 wm-bot2: Upgraded and rebooted host cloudcontrol1005.wikimedia.org - cookbook ran by andrew@bullseye === 2023-01-13 === * 17:36 wm-bot2: Restarting openstack services on cloudcontrol2005-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:36 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:36 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:35 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:35 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:35 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:33 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:33 wm-bot2: Restarting openstack services on cloudvirt2002-dev: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:33 wm-bot2: Restarting openstack services on cloudnet2005-dev: ['neutron-metadata-agent', 'neutron-dhcp-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:33 wm-bot2: Restarting openstack services on cloudvirt2003-dev: ['neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:33 wm-bot2: Restarting openstack services on cloudnet2006-dev: ['neutron-dhcp-agent', 'neutron-metadata-agent', 'neutron-linuxbridge-agent'] - cookbook ran by andrew@bullseye * 17:26 wm-bot2: Restarting openstack services on cloudvirt2001-dev: ['nova-compute'] - cookbook ran by andrew@bullseye * 17:26 wm-bot2: Restarting openstack services on cloudcontrol2004-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:25 wm-bot2: Restarting openstack services on cloudcontrol2001-dev: ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'] - cookbook ran by andrew@bullseye * 17:21 wm-bot2: Restarting openstack services: <nowiki>{</nowiki>'cloudcontrol2001-dev': ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'], 'cloudcontrol2004-dev': ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata'], 'cloudvirt2001-dev': ['nova-compute'], 'cloudvirt2002-dev': ['nova-compute'], 'cloudvirt2003-dev': ['nova-compute'], 'cloudcontrol2005-dev': ['nova-conductor', 'nova-scheduler', 'nova-api', 'nova-api-metadata * 10:41 arturo: restart backup_vm.service on cloudbackup1003/1004 to recover from a nova Unknown Error (HTTP 503) === 2023-01-12 === * 22:34 andrewbogott: updated the Bullseye base image with the upstream {{Gerrit|20221219}} build * 14:12 wm-bot2: Created project checkuser-beta-wiki with default quotas. ([[phab:T326740|T326740]]) - cookbook ran by arturo@nostromo === 2023-01-06 === * 18:42 wm-bot2: Safe reboot of cloudvirt2003-dev.codfw.wmnet finished successfully - cookbook ran by andrew@bullseye * 18:42 wm-bot2: unset cloudvirt2003-dev.codfw.wmnet maintenance (aggregates: ceph) - cookbook ran by andrew@bullseye * 18:39 wm-bot2: Drained cloudvirt2003-dev.codfw.wmnet - cookbook ran by andrew@bullseye * 18:35 wm-bot2: Set cloudvirt cloudvirt2003-dev.codfw.wmnet maintenance (downtime id: b50d9d3a-4f1d-4522-b86f-{{Gerrit|722fc9c55c87}}, use this to unset) - cookbook ran by andrew@bullseye * 18:35 wm-bot2: Draining cloudvirt2003-dev.codfw.wmnet - cookbook ran by andrew@bullseye * 18:35 wm-bot2: Safe rebooting cloudvirt2003-dev.codfw.wmnet - cookbook ran by andrew@bullseye * 18:35 wm-bot2: Draining cloudvirt2003-dev.cdofw.wmnet - cookbook ran by andrew@bullseye * 18:35 wm-bot2: Safe rebooting cloudvirt2003-dev.cdofw.wmnet - cookbook ran by andrew@bullseye * 00:14 wm-bot2: Upgraded and rebooted host cloudbackup1002-dev.eqiad.wmnet - cookbook ran by andrew@bullseye * 00:08 wm-bot2: Upgraded and rebooted host cloudbackup1001-dev.eqiad.wmnet - cookbook ran by andrew@bullseye * 00:01 wm-bot2: Upgraded and rebooted host cloudnet2006-dev.codfw.wmnet - cookbook ran by andrew@bullseye === 2023-01-05 === * 23:54 wm-bot2: Upgraded and rebooted host cloudnet2005-dev.codfw.wmnet - cookbook ran by andrew@bullseye * 23:38 wm-bot2: Upgraded and rebooted host cloudcontrol2005-dev.wikimedia.org - cookbook ran by andrew@bullseye * 23:25 wm-bot2: Upgraded and rebooted host cloudcontrol2004-dev.wikimedia.org - cookbook ran by andrew@bullseye * 23:13 wm-bot2: Upgraded and rebooted host cloudcontrol2001-dev.wikimedia.org - cookbook ran by andrew@bullseye * 22:18 wm-bot2: Upgraded and rebooted host cloudcontrol2001-dev.wikimedia.org - cookbook ran by andrew@bullseye * 22:10 andrewbogott: upgrading codfw1dev openstack to version 'zed' === 2023-01-04 === * 21:18 wm-bot2: Upgraded and rebooted host cloudservices1004.wikimedia.org - cookbook ran by andrew@bullseye * 21:11 wm-bot2: Upgraded and rebooted host cloudservices1005.wikimedia.org - cookbook ran by andrew@bullseye * 20:12 wm-bot2: Upgraded and rebooted host cloudservices1004.wikimedia.org - cookbook ran by andrew@bullseye * 20:04 wm-bot2: Upgraded and rebooted host cloudservices1005.wikimedia.org - cookbook ran by andrew@bullseye * 14:45 wm-bot2: Finished rebooting the nodes ['cloudcephmon2004-dev', 'cloudcephmon2005-dev', 'cloudcephmon2006-dev'] - cookbook ran by fran@wmf3169 * 14:44 wm-bot2: Finished rebooting node cloudcephmon2006-dev.codfw.wmnet - cookbook ran by fran@wmf3169 * 14:41 wm-bot2: Rebooting node cloudcephmon2006-dev.codfw.wmnet - cookbook ran by fran@wmf3169 * 14:41 wm-bot2: Finished rebooting node cloudcephmon2005-dev.codfw.wmnet - cookbook ran by fran@wmf3169 * 14:38 wm-bot2: Rebooting node cloudcephmon2005-dev.codfw.wmnet - cookbook ran by fran@wmf3169 * 14:37 wm-bot2: Finished rebooting node cloudcephmon2004-dev.codfw.wmnet - cookbook ran by fran@wmf3169 * 14:34 wm-bot2: Rebooting node cloudcephmon2004-dev.codfw.wmnet - cookbook ran by fran@wmf3169 * 14:34 wm-bot2: Rebooting the nodes cloudcephmon2004-dev,cloudcephmon2005-dev,cloudcephmon2006-dev - cookbook ran by fran@wmf3169 === 2023-01-03 === * 22:11 wm-bot2: Upgraded and rebooted host cloudservices2005-dev.wikimedia.org - cookbook ran by andrew@bullseye * 21:55 wm-bot2: Upgraded and rebooted host cloudservices2004-dev.wikimedia.org - cookbook ran by andrew@bullseye * 15:25 taavi: restart designate-sink everywhere to pick up wmf-sink changes === 2022-12-25 === * 14:21 taavi: register developer account 'instance-puppet-user-dev' to update the codfw1dev instance-puppet repo without access to the eqiad1 repo [[phab:T318504|T318504]] === 2022-12-22 === * 15:16 dcaro: added submit rights for JenkinsBot on all cloud/* gerrit repos === 2022-12-21 === * 04:59 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 04:59 wm-bot2: Finished rebooting node cloudcephosd1029.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 04:35 wm-bot2: Rebooting node cloudcephosd1029.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 04:35 wm-bot2: Finished rebooting node cloudcephosd1028.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 04:12 wm-bot2: Rebooting node cloudcephosd1028.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 04:12 wm-bot2: Finished rebooting node cloudcephosd1027.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 03:48 wm-bot2: Rebooting node cloudcephosd1027.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 03:48 wm-bot2: Finished rebooting node cloudcephosd1026.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 03:24 wm-bot2: Rebooting node cloudcephosd1026.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 03:24 wm-bot2: Finished rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 03:00 wm-bot2: Rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 03:00 wm-bot2: Finished rebooting node cloudcephosd1024.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 02:36 wm-bot2: Rebooting node cloudcephosd1024.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 02:36 wm-bot2: Finished rebooting node cloudcephosd1023.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 02:12 wm-bot2: Rebooting node cloudcephosd1023.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 02:12 wm-bot2: Finished rebooting node cloudcephosd1022.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 01:47 wm-bot2: Rebooting node cloudcephosd1022.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 01:47 wm-bot2: Finished rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 01:22 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 01:22 wm-bot2: Finished rebooting node cloudcephosd1020.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 00:58 wm-bot2: Rebooting node cloudcephosd1020.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 00:58 wm-bot2: Finished rebooting node cloudcephosd1019.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 00:34 wm-bot2: Rebooting node cloudcephosd1019.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 00:34 wm-bot2: Finished rebooting node cloudcephosd1018.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 00:09 wm-bot2: Rebooting node cloudcephosd1018.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 00:09 wm-bot2: Finished rebooting node cloudcephosd1017.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye === 2022-12-20 === * 23:45 wm-bot2: Rebooting node cloudcephosd1017.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 23:45 wm-bot2: Finished rebooting node cloudcephosd1016.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 23:21 wm-bot2: Rebooting node cloudcephosd1016.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 23:21 wm-bot2: Finished rebooting node cloudcephosd1015.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 22:56 wm-bot2: Rebooting node cloudcephosd1015.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 22:55 wm-bot2: Finished rebooting node cloudcephosd1014.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Rebooting node cloudcephosd1014.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 22:30 wm-bot2: Finished rebooting node cloudcephosd1013.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 22:08 andrewbogott: restarting openstack services with wmcs.openstack.restart_openstack due to miscellaneous trove failures * 22:06 wm-bot2: Rebooting node cloudcephosd1013.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 22:06 wm-bot2: Finished rebooting node cloudcephosd1012.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 21:40 wm-bot2: Rebooting node cloudcephosd1012.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 21:40 wm-bot2: Finished rebooting node cloudcephosd1011.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 21:15 wm-bot2: Rebooting node cloudcephosd1011.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 21:15 wm-bot2: Finished rebooting node cloudcephosd1010.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 20:50 wm-bot2: Rebooting node cloudcephosd1010.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 20:50 wm-bot2: Finished rebooting node cloudcephosd1009.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 20:26 wm-bot2: Rebooting node cloudcephosd1009.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 20:26 wm-bot2: Finished rebooting node cloudcephosd1008.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 20:01 wm-bot2: Rebooting node cloudcephosd1008.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 20:01 wm-bot2: Finished rebooting node cloudcephosd1007.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 19:35 wm-bot2: Rebooting node cloudcephosd1007.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 19:35 wm-bot2: Finished rebooting node cloudcephosd1006.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 19:10 wm-bot2: Rebooting node cloudcephosd1006.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 19:10 wm-bot2: Finished rebooting node cloudcephosd1005.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 18:45 wm-bot2: Rebooting node cloudcephosd1005.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 18:45 wm-bot2: Finished rebooting node cloudcephosd1004.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 18:20 wm-bot2: Rebooting node cloudcephosd1004.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 18:20 wm-bot2: Finished rebooting node cloudcephosd1003.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 17:56 wm-bot2: Rebooting node cloudcephosd1003.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 17:56 wm-bot2: Finished rebooting node cloudcephosd1002.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 17:33 wm-bot2: Rebooting node cloudcephosd1002.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 17:33 wm-bot2: Finished rebooting node cloudcephosd1001.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 17:07 wm-bot2: Rebooting node cloudcephosd1001.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 17:07 wm-bot2: Rebooting the nodes cloudcephosd1001,cloudcephosd1002,cloudcephosd1003,cloudcephosd1004,cloudcephosd1005,cloudcephosd1006,cloudcephosd1007,cloudcephosd1008,cloudcephosd1009,cloudcephosd1010,cloudcephosd1011,cloudcephosd1012,cloudcephosd1013,cloudcephosd1014,cloudcephosd1015,cloudcephosd1016,cloudcephosd1017,cloudcephosd1018,cloudcephosd1019,cloudcephosd1020,cloudcephosd1021,cloudcephosd1022,cloudcephosd1023,cloudcephosd1024,cl * 16:49 wm-bot2: Finished rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:48 wm-bot2: Finished rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:46 wm-bot2: Rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:46 wm-bot2: Finished rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:42 wm-bot2: Finished rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Rebooting node cloudcephosd1001.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:40 wm-bot2: Rebooting the nodes cloudcephosd1001,cloudcephosd1002,cloudcephosd1003,cloudcephosd1004,cloudcephosd1005,cloudcephosd1006,cloudcephosd1007,cloudcephosd1008,cloudcephosd1009,cloudcephosd1010,cloudcephosd1011,cloudcephosd1012,cloudcephosd1013,cloudcephosd1014,cloudcephosd1015,cloudcephosd1016,cloudcephosd1017,cloudcephosd1018,cloudcephosd1019,cloudcephosd1020,cloudcephosd1021,cloudcephosd1022,cloudcephosd1023,cloudcephosd1024,cl * 16:39 wm-bot2: Rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 16:39 wm-bot2: Rebooting the nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:41 wm-bot2: Rebooting node cloudcephosd1001.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:41 wm-bot2: Rebooting the nodes cloudcephosd1001,cloudcephosd1002,cloudcephosd1003,cloudcephosd1004,cloudcephosd1005,cloudcephosd1006,cloudcephosd1007,cloudcephosd1008,cloudcephosd1009,cloudcephosd1010,cloudcephosd1011,cloudcephosd1012,cloudcephosd1013,cloudcephosd1014,cloudcephosd1015,cloudcephosd1016,cloudcephosd1017,cloudcephosd1018,cloudcephosd1019,cloudcephosd1020,cloudcephosd1021,cloudcephosd1022,cloudcephosd1023,cloudcephosd1024,cl * 15:40 wm-bot2: Finished rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:40 wm-bot2: Finished rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:37 wm-bot2: Rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:37 wm-bot2: Finished rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:33 wm-bot2: Finished rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:30 wm-bot2: Rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:30 wm-bot2: Rebooting the nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:29 wm-bot2: Finished rebooting the nodes ['cloudcephmon2004-dev', 'cloudcephmon2005-dev', 'cloudcephmon2006-dev'] ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:29 wm-bot2: Finished rebooting node cloudcephmon2006-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:25 wm-bot2: Rebooting node cloudcephmon2006-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:25 wm-bot2: Finished rebooting node cloudcephmon2005-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:22 wm-bot2: Rebooting node cloudcephmon2005-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:22 wm-bot2: Finished rebooting node cloudcephmon2004-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:19 wm-bot2: Rebooting node cloudcephmon2004-dev.codfw.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:19 wm-bot2: Rebooting the nodes cloudcephmon2004-dev,cloudcephmon2005-dev,cloudcephmon2006-dev ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:18 wm-bot2: Finished rebooting the nodes ['cloudcephmon1001', 'cloudcephmon1002', 'cloudcephmon1003'] ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:18 wm-bot2: Finished rebooting node cloudcephmon1003.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:15 wm-bot2: Rebooting node cloudcephmon1003.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:15 wm-bot2: Finished rebooting node cloudcephmon1002.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:11 wm-bot2: Rebooting node cloudcephmon1002.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:11 wm-bot2: Finished rebooting node cloudcephmon1001.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:08 wm-bot2: Rebooting node cloudcephmon1001.eqiad.wmnet ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye * 15:08 wm-bot2: Rebooting the nodes cloudcephmon1001,cloudcephmon1002,cloudcephmon1003 ([[phab:T325132|T325132]]) - cookbook ran by andrew@bullseye === 2022-12-17 === * 07:50 taavi: deleted project packagist-mirror per https://wikitech.wikimedia.org/wiki/News/Cloud_VPS_2022_Purge#packagist-mirror === 2022-12-16 === * 19:36 volans: restarted sshd twice on bastion-restricted-eqiad1-02 to debug SSH connections for [[phab:T319401|T319401]] * 08:46 dcaro: restart designate-sink on both cloudservice hosts ([[phab:T322279|T322279]]) * 08:45 dcaro: restart designate-sink on both cloudservice hosts === 2022-12-07 === * 22:07 andrewbogott: systemctl restart libvirt-guests.service on cloudvirt1019 to get ceph/rbd working on VMS on this hypervisor === 2022-12-03 === * 19:24 taavi: restart designate-sink on both cloudservices hosts === 2022-11-30 === * 20:03 andrewbogott: changing all rabbitmq queues to quorum queues. Will be noisy! [[phab:T318816|T318816]] * 02:54 wm-bot2: Upgraded and rebooted host cloudbackup2002.codfw.wmnet - cookbook ran by andrew@bullseye === 2022-11-28 === * 13:00 wm-bot2: unset cloudvirt1043.eqiad.wmnet maintenance (aggregates: ceph) - cookbook ran by arturo@nostromo * 10:28 wm-bot2: Drained cloudvirt1043.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 10:19 wm-bot2: Set cloudvirt cloudvirt1043.eqiad.wmnet maintenance (downtime id: bb94dd24-fef9-4c9c-8f79-{{Gerrit|b6e15023ce69}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 10:18 wm-bot2: Draining cloudvirt1043.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo === 2022-11-25 === * 10:54 wm-bot2: deleted VM canary2001-dev-2 from cloudvirt2001-dev - cookbook ran by arturo@nostromo * 10:54 wm-bot2: created VM canary2001-dev-3 in cloudvirt2001-dev - cookbook ran by arturo@nostromo * 10:53 wm-bot2: Created new flavor: g3.cores1.ram1.disk20 (id:5b2ca632-2ea0-4007-9b40-{{Gerrit|4f84f8e2428b}}) - cookbook ran by arturo@nostromo * 10:46 wm-bot2: Created new flavor: g3.cores1.ram1.disk20.admin (id:fecbd56d-0969-45f3-80fd-{{Gerrit|2b463a5b6270}}) - cookbook ran by arturo@nostromo === 2022-11-24 === * 16:53 wm-bot2: deleted VM canary2001-dev-1 from cloudvirt2001-dev - cookbook ran by arturo@nostromo * 16:53 wm-bot2: created VM canary2001-dev-2 in cloudvirt2001-dev - cookbook ran by arturo@nostromo * 16:42 wm-bot2: created VM canary2003-dev-1 in cloudvirt2003-dev - cookbook ran by arturo@nostromo * 16:42 wm-bot2: created VM canary2002-dev-1 in cloudvirt2002-dev - cookbook ran by arturo@nostromo * 16:42 wm-bot2: created VM canary2001-dev-1 in cloudvirt2001-dev - cookbook ran by arturo@nostromo * 16:36 wm-bot2: Created new flavor: cloudvirt-canary-ceph (id:0d06701a-2845-4298-b2b4-{{Gerrit|fabf8b1ddcbb}}) - cookbook ran by arturo@nostromo * 13:03 wm-bot2: unset cloudvirt1044.eqiad.wmnet maintenance (aggregates: ceph) - cookbook ran by arturo@nostromo * 11:51 wm-bot2: Drained cloudvirt1044.eqiad.wmnet - cookbook ran by arturo@nostromo * 11:37 wm-bot2: Set cloudvirt cloudvirt1044.eqiad.wmnet maintenance (downtime id: 10076ecc-0f94-4d56-9bbd-{{Gerrit|bebb48bdc126}}, use this to unset) - cookbook ran by arturo@nostromo * 11:35 wm-bot2: Draining cloudvirt1044.eqiad.wmnet - cookbook ran by arturo@nostromo * 10:16 dcaro: removed ip6 dns name entry from nb for coluddb* ([[phab:T323550|T323550]]) * 09:53 dcaro: removed ip6 dns entry from nb for coluddb1013 ([[phab:T323550|T323550]]) === 2022-11-23 === * 15:00 wm-bot2: unset cloudvirt1045.eqiad.wmnet maintenance (aggregates: ceph) - cookbook ran by arturo@nostromo * 13:46 wm-bot2: Drained cloudvirt1045.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 13:35 wm-bot2: Set cloudvirt cloudvirt1045.eqiad.wmnet maintenance (downtime id: 2386b468-0f21-4ecb-91e2-{{Gerrit|e19ace66881d}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 13:34 wm-bot2: Draining cloudvirt1045.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 13:20 wm-bot2: unset cloudvirt1046.eqiad.wmnet maintenance (aggregates: ceph) - cookbook ran by arturo@nostromo * 12:17 wm-bot2: Drained cloudvirt1046.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 12:04 wm-bot2: Set cloudvirt cloudvirt1046.eqiad.wmnet maintenance (downtime id: 6291d38e-c04c-4aa0-88da-{{Gerrit|2b329874a9b9}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 12:03 wm-bot2: Draining cloudvirt1046.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 12:02 wm-bot2: unset cloudvirt1047.eqiad.wmnet maintenance (aggregates: ceph) - cookbook ran by arturo@nostromo * 11:32 arturo: [codfw1dev] created project cloudvirt-canary * 10:13 wm-bot2: Drained cloudvirt1047.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 10:01 wm-bot2: Set cloudvirt cloudvirt1047.eqiad.wmnet maintenance (downtime id: 2c7ccc17-2be3-427d-aaec-{{Gerrit|57fadca0de5b}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 10:00 wm-bot2: Draining cloudvirt1047.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo === 2022-11-22 === * 13:26 wm-bot2: unset cloudvirt1048.eqiad.wmnet maintenance (aggregates: ceph) - cookbook ran by arturo@nostromo * 12:15 wm-bot2: unset cloudvirt1049.eqiad.wmnet maintenance (aggregates: ceph) - cookbook ran by arturo@nostromo * 10:58 wm-bot2: Drained cloudvirt1049.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 10:37 wm-bot2: Set cloudvirt cloudvirt1049.eqiad.wmnet maintenance (downtime id: 3ae0a20d-3cf0-4eba-bea6-{{Gerrit|45aa61d8ad00}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 10:36 wm-bot2: Draining cloudvirt1049.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 10:21 wm-bot2: Unset cloudvirt cloudvirt1050.eqiad.wmnet maintenance - cookbook ran by arturo@nostromo === 2022-11-21 === * 16:19 wm-bot2: Drained cloudvirt1050.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 16:11 wm-bot2: Set cloudvirt cloudvirt1050.eqiad.wmnet maintenance (downtime id: 0b154a93-a9d3-4ac7-bcfc-{{Gerrit|b67c49abe97b}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 16:09 wm-bot2: Draining cloudvirt1050.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 16:07 wm-bot2: Unset cloudvirt cloudvirt1051.eqiad.wmnet maintenance - cookbook ran by arturo@nostromo * 15:18 wm-bot2: Drained cloudvirt1051.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 15:02 wm-bot2: Set cloudvirt cloudvirt1051.eqiad.wmnet maintenance (downtime id: 1de1174b-ec46-47a3-911c-{{Gerrit|b5808ce37028}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 15:00 wm-bot2: Draining cloudvirt1051.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 14:57 wm-bot2: Unset cloudvirt cloudvirt1052.eqiad.wmnet maintenance - cookbook ran by arturo@nostromo * 12:53 wm-bot2: Drained cloudvirt1052.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 12:30 wm-bot2: Set cloudvirt cloudvirt1052.eqiad.wmnet maintenance (downtime id: 30db8a9b-08db-456f-8106-{{Gerrit|53188ff5f989}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 12:29 wm-bot2: Draining cloudvirt1052.eqiad.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 12:15 wm-bot2: Unset cloudvirt cloudvirt1053.eqiad.wmnet maintenance - cookbook ran by arturo@nostromo * 10:13 arturo: drained cloudvirt1053 in preparation for reimage (was spare anyway) === 2022-11-18 === * 13:37 arturo: [codfw1dev] reimaged cloudvirt2001-dev and cloudvirt2002-dev * 11:36 wm-bot2: Set cloudvirt cloudvirt2001-dev.codfw.wmnet maintenance (downtime id: 2fb9df34-fc2d-45b9-b21f-{{Gerrit|c0ec09008b92}}, use this to unset) - cookbook ran by arturo@nostromo * 11:36 wm-bot2: Draining cloudvirt2001-dev.codfw.wmnet - cookbook ran by arturo@nostromo === 2022-11-16 === * 20:07 wm-bot2: Upgraded and rebooted host cloudcontrol2004-dev.wikimedia.org - cookbook ran by andrew@bullseye * 19:56 wm-bot2: Upgraded and rebooted host cloudservices2004-dev.wikimedia.org - cookbook ran by andrew@bullseye * 19:16 wm-bot2: Upgraded and rebooted host cloudcontrol1007.wikimedia.org - cookbook ran by andrew@bullseye * 18:59 wm-bot2: Upgraded and rebooted host cloudcontrol1007.wikimedia.org - cookbook ran by andrew@bullseye * 18:51 wm-bot2: Upgraded and rebooted host cloudcontrol2005-dev.wikimedia.org - cookbook ran by andrew@bullseye * 12:52 arturo: failovered cloudgw1002 into cloudgw1001 for reimage, IRC bots were briefly disconnected === 2022-11-14 === * 20:22 wm-bot2: Upgraded and rebooted host cloudnet1005.eqiad.wmnet - cookbook ran by andrew@bullseye * 20:07 wm-bot2: Upgraded and rebooted host cloudcontrol1007.wikimedia.org - cookbook ran by andrew@bullseye * 19:55 wm-bot2: Upgraded and rebooted host cloudcontrol1006.wikimedia.org - cookbook ran by andrew@bullseye * 19:42 wm-bot2: Upgraded and rebooted host cloudcontrol1005.wikimedia.org - cookbook ran by andrew@bullseye * 19:28 andrewbogott: beginning OpenStack upgrade in eqiad1 -- [[phab:T305828|T305828]] * 19:22 wm-bot2: Upgraded and rebooted host cloudbackup1002-dev.eqiad.wmnet - cookbook ran by andrew@bullseye * 19:14 wm-bot2: Upgraded and rebooted host cloudbackup1001-dev.eqiad.wmnet - cookbook ran by andrew@bullseye * 19:08 wm-bot2: Upgraded and rebooted host cloudcontrol2001-dev.wikimedia.org - cookbook ran by andrew@bullseye * 12:52 arturo: cleanup old network vlan interface names from /etc/network/interfaces in cloudnet1005/1006 === 2022-11-11 === * 13:24 wm-bot2: Set cloudvirt cloudvirt2003-dev.codfw.wmnet maintenance (downtime id: edad3915-b7c6-4b23-bb9c-{{Gerrit|ab13b04a41c5}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 13:23 wm-bot2: Draining cloudvirt2003-dev.codfw.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 13:17 wm-bot2: Set cloudvirt cloudvirt2003-dev.codfw.wmnet maintenance (downtime id: a09a9868-8aae-4dc0-8510-{{Gerrit|a3923b703060}}, use this to unset) ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 13:16 wm-bot2: Draining cloudvirt2003-dev.codfw.wmnet ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo === 2022-11-10 === * 16:19 wm-bot2: Set cloudvirt 'cloudvirt2002-dev.codfw.wmnet' maintenance (downtime id: 346013ec-ce4e-497e-ad65-{{Gerrit|2d215b14998c}}, use this to unset). ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 16:18 wm-bot2: Draining 'cloudvirt2002-dev.codfw.wmnet'. ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo * 16:13 wm-bot2: Set cloudvirt 'cloudvirt2002-dev.codfw.wmnet' maintenance (downtime id: a40a8487-22ee-4bc6-bbe4-{{Gerrit|694a615e3bf5}}, use this to unset). ([[phab:T319184|T319184]]) - cookbook ran by arturo@nostromo === 2022-11-08 === * 11:17 taavi: backfilling security groups for metricsinfra access on all projects [[phab:T288108|T288108]] === 2022-11-07 === * 21:01 wm-bot2: Upgraded and rebooted host cloudservices1004.wikimedia.org ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 20:50 wm-bot2: Upgraded and rebooted host cloudservices1005.wikimedia.org ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 20:40 andrewbogott: upgrading eqiad1 designate to version 'yoga' === 2022-11-04 === * 17:57 andrewbogott: removing cinderv2 API endpoints from keystone catalog; this is deprecated and removed in Yoga. prep for [[phab:T305828|T305828]] === 2022-11-03 === * 19:24 wm-bot2: Upgraded and rebooted host cloudbackup1002-dev.eqiad.wmnet ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 19:19 wm-bot2: Upgraded and rebooted host cloudbackup1001-dev.eqiad.wmnet ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 00:08 wm-bot2: Upgraded and rebooted host cloudcontrol2004-dev.wikimedia.org ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye === 2022-11-02 === * 23:39 wm-bot2: Upgraded and rebooted host cloudbackup1001-dev.eqiad.wmnet ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 23:38 wm-bot2: Upgraded and rebooted host cloudnet2006-dev.codfw.wmnet ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 23:29 wm-bot2: Upgraded and rebooted host cloudnet2005-dev.codfw.wmnet ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 23:13 wm-bot2: Upgraded and rebooted host cloudcontrol2004-dev.wikimedia.org ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 23:00 wm-bot2: Upgraded and rebooted host cloudcontrol2001-dev.wikimedia.org ([[phab:T305828|T305828]]) - cookbook ran by andrew@bullseye * 22:54 wm-bot2: Upgraded and rebooted host cloudcontrol2005-dev.wikimedia.org - cookbook ran by andrew@bullseye * 20:29 wm-bot2: Upgraded and rebooted host cloudcontrol2001-dev.wikimedia.org - cookbook ran by andrew@bullseye === 2022-10-31 === * 13:09 arturo: restart keepalived on all 4 cloudgw servers to run them with `-D` in /etc/default/keepalived to further debug [[phab:T320975|T320975]] === 2022-10-26 === * 16:18 wm-bot2: Created new flavor: g3.cores1.ram1.disk20 (id:bf48880d-0c1b-4c2a-8e8b-{{Gerrit|778d28b16561}}) ([[phab:T319446|T319446]]) - cookbook ran by dcaro@vulcanus * 09:34 taavi: running wmcs-puppetcertleaks in delete mode * 09:09 taavi: running wmcs-novastats-dnsleaks in delete mode === 2022-10-25 === * 16:03 arturo: [codfw1dev] [[phab:T321220|T321220]] root@cloudcontrol2001-dev:~# openstack subnet create magnum --no-dhcp --network 57017d7c-3817-429a-8aa3-{{Gerrit|b028de82cdcc}} --ip-version 4 --gateway auto --subnet-range 192.168.0.0/24 * 14:38 arturo: [codfw1dev] restart neutron-l3-agent in cloudnet2005-dev, it was dead after rabbit connectivity problems === 2022-10-24 === * 18:50 wm-bot2: Rebooting node cloudcephmon1002.eqiad.wmnet - cookbook ran by andrew@bullseye * 18:50 wm-bot2: Finished rebooting node cloudcephmon1001.eqiad.wmnet - cookbook ran by andrew@bullseye * 18:47 wm-bot2: Rebooting node cloudcephmon1001.eqiad.wmnet - cookbook ran by andrew@bullseye * 18:47 wm-bot2: Rebooting the nodes cloudcephmon1001,cloudcephmon1002,cloudcephmon1003 - cookbook ran by andrew@bullseye * 18:38 wm-bot2: Rebooting the nodes cloudcephmon1001,cloudcephmon1002,cloudcephmon1003 - cookbook ran by andrew@bullseye * 18:21 wm-bot2: Rebooting node cloudcephosd1001.eqiad.wmnet - cookbook ran by andrew@bullseye * 18:21 wm-bot2: Rebooting the nodes cloudcephosd1001,cloudcephosd1002,cloudcephosd1003,cloudcephosd1004,cloudcephosd1005,cloudcephosd1006,cloudcephosd1007,cloudcephosd1008,cloudcephosd1009,cloudcephosd1010,cloudcephosd1011,cloudcephosd1012,cloudcephosd1013,cloudcephosd1014,cloudcephosd1015,cloudcephosd1016,cloudcephosd1017,cloudcephosd1018,cloudcephosd1019,cloudcephosd1020,cloudcephosd1021,cloudcephosd1022,cloudcephosd1023,cloudcephosd1024,cl === 2022-10-20 === * 23:23 wm-bot2: Safe reboot of 'cloudvirt1021.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 23:23 wm-bot2: Unset cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 23:20 wm-bot2: Safe reboot of 'cloudvirt1022.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 23:20 wm-bot2: Unset cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 23:19 wm-bot2: Drained 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 23:16 wm-bot2: Drained 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 23:07 wm-bot2: Safe reboot of 'cloudvirt1024.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 23:07 wm-bot2: Unset cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 23:06 wm-bot2: Safe reboot of 'cloudvirt1025.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 23:06 wm-bot2: Unset cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 23:03 wm-bot2: Drained 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 23:03 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: 101c41d1-d65d-4088-b6d2-{{Gerrit|eac859e45ef8}}, use this to unset). - cookbook ran by andrew@bullseye * 23:03 wm-bot2: Drained 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 23:03 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 23:03 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 23:03 wm-bot2: Safe reboot of 'cloudvirt1026.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 23:01 wm-bot2: Unset cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 22:57 wm-bot2: Drained 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:51 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: 7197ce34-cc57-4677-a1a9-{{Gerrit|05e20cd0dd80}}, use this to unset). - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Safe reboot of 'cloudvirt1017.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 22:50 wm-bot2: Unset cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 22:46 wm-bot2: Drained 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:38 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 68e4cc1d-9def-444e-84a7-{{Gerrit|21b0b5adfa72}}, use this to unset). - cookbook ran by andrew@bullseye * 22:37 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 73388c1e-369b-40af-a2c1-{{Gerrit|96936128c324}}, use this to unset). - cookbook ran by andrew@bullseye * 22:37 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance (downtime id: 12feeca2-ea59-48e9-9235-{{Gerrit|1cecf5b384cd}}, use this to unset). - cookbook ran by andrew@bullseye * 22:37 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:37 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:37 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:37 wm-bot2: Safe reboot of 'cloudvirt1029.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 22:37 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:36 wm-bot2: Unset cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 22:36 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:36 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:35 wm-bot2: Safe reboot of 'cloudvirt1030.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 22:35 wm-bot2: Unset cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 22:35 wm-bot2: Safe reboot of 'cloudvirt1027.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 22:35 wm-bot2: Unset cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 22:34 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: 1b34587f-2770-4d2e-bb31-{{Gerrit|c8ad14632d39}}, use this to unset). - cookbook ran by andrew@bullseye * 22:34 wm-bot2: Drained 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:34 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:33 wm-bot2: Drained 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:32 wm-bot2: Drained 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:28 wm-bot2: Safe reboot of 'cloudvirt1032.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 22:28 wm-bot2: Unset cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 22:24 wm-bot2: Drained 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:13 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance (downtime id: e75b00eb-7f58-4821-8139-{{Gerrit|3dfc6e97a92a}}, use this to unset). - cookbook ran by andrew@bullseye * 22:12 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: 9151511e-bcc0-4367-a976-{{Gerrit|7cff060308aa}}, use this to unset). - cookbook ran by andrew@bullseye * 22:12 wm-bot2: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance (downtime id: 241c0bd9-c3cf-40e4-95dc-{{Gerrit|9b51d9823fe9}}, use this to unset). - cookbook ran by andrew@bullseye * 22:12 wm-bot2: Set cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance (downtime id: 7a53c274-1d5f-4c94-9a0d-{{Gerrit|721c3e8f7239}}, use this to unset). - cookbook ran by andrew@bullseye * 22:12 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:12 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:11 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:11 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:11 wm-bot2: Draining 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:11 wm-bot2: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:11 wm-bot2: Draining 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:11 wm-bot2: Safe rebooting 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 22:10 wm-bot2: Safe reboot of 'cloudvirt1031.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 22:10 wm-bot2: Unset cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 22:06 wm-bot2: Drained 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:58 wm-bot2: Safe reboot of 'cloudvirt1034.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 21:58 wm-bot2: Unset cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 21:58 wm-bot2: Safe reboot of 'cloudvirt1035.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 21:58 wm-bot2: Unset cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 21:55 wm-bot2: Safe reboot of 'cloudvirt1033.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 21:55 wm-bot2: Unset cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 21:54 wm-bot2: Drained 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:54 wm-bot2: Drained 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:51 wm-bot2: Drained 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:38 wm-bot2: Set cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance (downtime id: 0b75b44c-1efe-4edf-8f5e-{{Gerrit|42a67e8d3b13}}, use this to unset). - cookbook ran by andrew@bullseye * 21:38 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: aeb9c3a6-961a-4873-85d4-{{Gerrit|929248aebb8b}}, use this to unset). - cookbook ran by andrew@bullseye * 21:38 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: eb14cc04-3293-4cf8-a46f-{{Gerrit|1f87d8b0bcc4}}, use this to unset). - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: 84a3cd99-2643-495a-86b6-{{Gerrit|7ae80b11a30d}}, use this to unset). - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Draining 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Safe rebooting 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Draining 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:37 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Draining 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:36 wm-bot2: Safe rebooting 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:32 wm-bot2: Safe reboot of 'cloudvirt1036.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 21:32 wm-bot2: Unset cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 21:28 wm-bot2: Drained 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:28 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance (downtime id: 01a0b69c-2e27-4331-9635-{{Gerrit|403a668dac29}}, use this to unset). - cookbook ran by andrew@bullseye * 21:27 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:27 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:23 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance (downtime id: ea29f65f-6d86-4568-8db9-{{Gerrit|a85faa827447}}, use this to unset). - cookbook ran by andrew@bullseye * 21:22 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:22 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:19 wm-bot2: Safe reboot of 'cloudvirt1038.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 21:19 wm-bot2: Unset cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 21:19 wm-bot2: Safe reboot of 'cloudvirt1037.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 21:19 wm-bot2: Unset cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 21:18 wm-bot2: Safe reboot of 'cloudvirt1039.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 21:18 wm-bot2: Unset cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 21:16 wm-bot2: Drained 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:15 wm-bot2: Drained 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:14 wm-bot2: Drained 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:01 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance (downtime id: d6b6bb4d-c7c0-4bd2-9a02-{{Gerrit|dee9a5ec6a3e}}, use this to unset). - cookbook ran by andrew@bullseye * 21:01 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance (downtime id: 4e109149-7d67-403f-8b21-{{Gerrit|f829235ea491}}, use this to unset). - cookbook ran by andrew@bullseye * 21:01 wm-bot2: Set cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance (downtime id: 341f4b8c-6d6e-4190-9f2d-{{Gerrit|a1a1483abadd}}, use this to unset). - cookbook ran by andrew@bullseye * 21:01 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: f76e8f14-920e-4d51-a44e-{{Gerrit|321a55dcb0c5}}, use this to unset). - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Draining 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 21:00 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:56 wm-bot2: Safe reboot of 'cloudvirt1042.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 20:56 wm-bot2: Unset cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 20:54 wm-bot2: Unset cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 20:35 wm-bot2: Drained 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:34 wm-bot2: Drained 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:32 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance (downtime id: 0b0d8090-4d31-4d43-8545-{{Gerrit|c25209e2ef58}}, use this to unset). - cookbook ran by andrew@bullseye * 20:31 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:31 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:21 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance (downtime id: 3d928554-a79c-419a-bcc4-{{Gerrit|c8d63791d8e7}}, use this to unset). - cookbook ran by andrew@bullseye * 20:21 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance (downtime id: 3d3c8aa5-45f7-429e-a9d7-{{Gerrit|b5181b29dc13}}, use this to unset). - cookbook ran by andrew@bullseye * 20:21 wm-bot2: Set cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance (downtime id: b9027020-8a90-4c40-b0e6-{{Gerrit|6c8767f87917}}, use this to unset). - cookbook ran by andrew@bullseye * 20:21 wm-bot2: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance (downtime id: f0d60fab-64e9-4da2-826c-{{Gerrit|229d31d0fbc2}}, use this to unset). - cookbook ran by andrew@bullseye * 20:21 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:21 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Draining 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Safe rebooting 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Draining 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Safe rebooting 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Safe reboot of 'cloudvirt1044.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 20:20 wm-bot2: Unset cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 20:19 wm-bot2: Safe reboot of 'cloudvirt1047.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 20:19 wm-bot2: Unset cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 20:18 wm-bot2: Safe reboot of 'cloudvirt1046.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 20:18 wm-bot2: Unset cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 20:16 wm-bot2: Drained 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:15 wm-bot2: Drained 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:15 wm-bot2: Safe reboot of 'cloudvirt1045.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 20:15 wm-bot2: Unset cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 20:14 wm-bot2: Drained 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 20:11 wm-bot2: Drained 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:55 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: b65da7e2-9fd2-4a1e-8dd4-{{Gerrit|88ca65936ae4}}, use this to unset). - cookbook ran by andrew@bullseye * 19:55 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance (downtime id: cbbf114c-9a4c-4dd2-9b58-{{Gerrit|50da8bd896ca}}, use this to unset). - cookbook ran by andrew@bullseye * 19:55 wm-bot2: Set cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance (downtime id: 5969beaf-72bf-4af9-ae4d-{{Gerrit|8e5c3331e78e}}, use this to unset). - cookbook ran by andrew@bullseye * 19:55 wm-bot2: Set cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance (downtime id: b8f7389d-79b4-411d-b3d5-{{Gerrit|ec8f92eba101}}, use this to unset). - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Draining 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Safe rebooting 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Draining 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:54 wm-bot2: Safe rebooting 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:53 wm-bot2: Safe reboot of 'cloudvirt1051.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 19:53 wm-bot2: Unset cloudvirt 'cloudvirt1051.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 19:51 wm-bot2: Safe reboot of 'cloudvirt1049.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 19:51 wm-bot2: Unset cloudvirt 'cloudvirt1049.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 19:50 wm-bot2: Safe reboot of 'cloudvirt1048.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 19:50 wm-bot2: Unset cloudvirt 'cloudvirt1048.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 19:49 wm-bot2: Drained 'cloudvirt1051.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Set cloudvirt 'cloudvirt1050.eqiad.wmnet' maintenance (downtime id: c5dbfa7b-72fc-4156-8257-{{Gerrit|af224a725b78}}, use this to unset). - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Draining 'cloudvirt1050.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:36 wm-bot2: Safe rebooting 'cloudvirt1050.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Draining 'cloudvirt1050.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:33 wm-bot2: Safe rebooting 'cloudvirt1050.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:30 wm-bot2: Set cloudvirt 'cloudvirt1048.eqiad.wmnet' maintenance (downtime id: c1a3e92c-cb04-46e4-ac5d-{{Gerrit|784c79601b05}}, use this to unset). - cookbook ran by andrew@bullseye * 19:30 wm-bot2: Set cloudvirt 'cloudvirt1049.eqiad.wmnet' maintenance (downtime id: 62fdc4ee-c8d5-4892-aeb9-{{Gerrit|a681cdcbc84b}}, use this to unset). - cookbook ran by andrew@bullseye * 19:29 wm-bot2: Set cloudvirt 'cloudvirt1050.eqiad.wmnet' maintenance (downtime id: 9ec407cd-5c00-407a-a57d-{{Gerrit|794f6c68f947}}, use this to unset). - cookbook ran by andrew@bullseye * 19:29 wm-bot2: Set cloudvirt 'cloudvirt1051.eqiad.wmnet' maintenance (downtime id: 712683ea-d58f-484b-8efb-{{Gerrit|89d1c21aa0e3}}, use this to unset). - cookbook ran by andrew@bullseye * 19:29 wm-bot2: Draining 'cloudvirt1048.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:29 wm-bot2: Safe rebooting 'cloudvirt1048.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:29 wm-bot2: Draining 'cloudvirt1049.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:29 wm-bot2: Safe rebooting 'cloudvirt1049.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:28 wm-bot2: Draining 'cloudvirt1050.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:28 wm-bot2: Safe rebooting 'cloudvirt1050.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:28 wm-bot2: Draining 'cloudvirt1051.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:28 wm-bot2: Safe rebooting 'cloudvirt1051.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:25 wm-bot2: Safe reboot of 'cloudvirt1052.eqiad.wmnet' finished successfully. - cookbook ran by andrew@bullseye * 19:25 wm-bot2: Unset cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance. - cookbook ran by andrew@bullseye * 19:21 wm-bot2: Drained 'cloudvirt1052.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:04 wm-bot2: Set cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance (downtime id: 18fae8d8-7353-4f67-90d7-{{Gerrit|8df9b3fb1ccb}}, use this to unset). - cookbook ran by andrew@bullseye * 19:03 wm-bot2: Draining 'cloudvirt1052.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:03 wm-bot2: Safe rebooting 'cloudvirt1052.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:01 wm-bot2: Set cloudvirt 'cloudvirt1053.eqiad.wmnet' maintenance (downtime id: 3d3cffa3-abc6-4901-83d9-{{Gerrit|3bcc4b02fd3c}}, use this to unset). - cookbook ran by andrew@bullseye * 19:00 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 19:00 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:54 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:54 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:52 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:52 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:51 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:51 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:51 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:51 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:51 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:50 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:50 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:50 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:46 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye * 18:46 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. - cookbook ran by andrew@bullseye === 2022-10-15 === * 17:38 taavi: taavi@cloudweb1003 ~ $ mwscript extensions/OATHAuth/maintenance/disableOATHAuthForUser.php --wiki=labswiki Slevinski # [[phab:T320867|T320867]] === 2022-10-13 === * 12:19 wm-bot2: OSDs (['cloudcephosd1027', 'cloudcephosd1028', 'cloudcephosd1029', 'cloudcephosd1030', 'cloudcephosd1031', 'cloudcephosd1032', 'cloudcephosd1033', 'cloudcephosd1034']) upgraded successfully B-) ([[phab:T309786|T309786]]) - cookbook ran by dcaro@vulcanus * 11:31 wm-bot2: Upgrading OSDs and rebooting the nodes ['cloudcephosd1027', 'cloudcephosd1028', 'cloudcephosd1029', 'cloudcephosd1030', 'cloudcephosd1031', 'cloudcephosd1032', 'cloudcephosd1033', 'cloudcephosd1034'] ([[phab:T309786|T309786]]) - cookbook ran by dcaro@vulcanus * 11:30 wm-bot2: OSDs (['cloudcephosd1025', 'cloudcephosd1026']) upgraded successfully B-) ([[phab:T309786|T309786]]) - cookbook ran by dcaro@vulcanus === 2022-10-10 === * 14:01 dcaro: test2 * 01:22 andrewbogott: restarting designate-sink on cloudservices100[45], possible example of [[phab:T316614|T316614]] === 2022-10-09 === * 12:04 taavi: taavi@cloudweb1003 ~ $ mwscript extensions/OATHAuth/maintenance/disableOATHAuthForUser.php --wiki=labswiki DatGuy # [[phab:T320301|T320301]] === 2022-10-07 === * 13:40 andrewbogott: dhinus is resetting rabbitmq cluster in an attempt to resolve a suspected (by Andrew) split-brain * 11:33 arturo: rabbitmq-server.service @ cloudrabbit1002 is again up and running ([[phab:T320232|T320232]]) * 10:24 arturo: stopping rabbitmq-server.service @ cloudrabbit1002 ([[phab:T320232|T320232]]) * 10:19 arturo: restarting nova-conductor in all 3 cloudcontrols ([[phab:T320232|T320232]]) * 09:45 arturo: restarting rabbitmq-server.service @ cloudrabbit1002 ([[phab:T320232|T320232]]) === 2022-10-06 === * 15:55 arturo: cloudnet1005 & cloudnet1006 now in service. Secom cloudnet1003 & cloudnet1004. Drop neutron agents, etc. ([[phab:T316284|T316284]]) * 11:54 arturo: rebooting cloudnet1005/1006 to see if they have the right network config ([[phab:T316284|T316284]]) * 11:50 arturo: set neutron l3 agents on cloudnet1005/1006 as down `root@cloudcontrol1005:~# neutron agent-update --admin-state-down <uuid>` ([[phab:T316284|T316284]]) * 11:40 arturo: [codfw1dev] rebooting both network nodes to test https://gerrit.wikimedia.org/r/c/operations/puppet/+/839492 * 10:14 arturo: [codfw1dev] restart neutron-l3-agent on cloudnet2006-dev, it was dead === 2022-10-05 === * 14:40 wm-bot2: Adding OSD cloudcephosd1021.eqiad.wmnet... (1/1) ([[phab:T319418|T319418]]) - cookbook ran by fran@wmf3169 * 14:40 wm-bot2: Adding new OSDs ['cloudcephosd1021.eqiad.wmnet'] to the cluster ([[phab:T319418|T319418]]) - cookbook ran by fran@wmf3169 * 14:28 arturo: adding cloudinstances2b-gw router to l3 agents on cloudnet1005/1006 ([[phab:T316284|T316284]]) * 13:11 wm-bot2: Added 1 new OSDs ['cloudcephosd1034.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 13:11 wm-bot2: Added OSD cloudcephosd1034.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 13:02 wm-bot2: Finished rebooting node cloudcephosd1034.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 12:58 wm-bot2: Rebooting node cloudcephosd1034.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 12:58 wm-bot2: Adding OSD cloudcephosd1034.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 12:58 wm-bot2: Adding new OSDs ['cloudcephosd1034.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 === 2022-10-04 === * 16:40 wm-bot2: Added 1 new OSDs ['cloudcephosd1033.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 16:40 wm-bot2: Added OSD cloudcephosd1033.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 14:34 wm-bot2: Finished rebooting node cloudcephosd1033.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 14:30 wm-bot2: Rebooting node cloudcephosd1033.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 14:30 wm-bot2: Adding OSD cloudcephosd1033.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 14:30 wm-bot2: Adding new OSDs ['cloudcephosd1033.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 14:20 wm-bot2: Finished rebooting node cloudcephosd1033.eqiad.wmnet - cookbook ran by fran@wmf3169 * 14:17 wm-bot2: Rebooting node cloudcephosd1033.eqiad.wmnet - cookbook ran by fran@wmf3169 * 14:16 wm-bot2: Adding OSD cloudcephosd1033.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 14:16 wm-bot2: Adding new OSDs ['cloudcephosd1033.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 10:59 wm-bot2: Finished rebooting node cloudcephosd1033.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:56 wm-bot2: Rebooting node cloudcephosd1033.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:55 wm-bot2: Adding OSD cloudcephosd1033.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 10:55 wm-bot2: Adding new OSDs ['cloudcephosd1033.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 === 2022-09-30 === * 14:52 wm-bot2: Added 1 new OSDs ['cloudcephosd1031.eqiad.wmnet'] - cookbook ran by fran@wmf3169 * 14:52 wm-bot2: Added OSD cloudcephosd1031.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 14:48 wm-bot2: Adding OSD cloudcephosd1031.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 14:48 wm-bot2: Adding new OSDs ['cloudcephosd1031.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 14:16 wm-bot2: Drained 'cloudvirt1023.eqiad.wmnet'. ([[phab:T319025|T319025]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: 64eac5c6-4b1d-4269-98fd-{{Gerrit|8e5bed42ce40}}, use this to unset). ([[phab:T319025|T319025]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. ([[phab:T319025|T319025]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. ([[phab:T319025|T319025]]) - cookbook ran by andrew@buster * 14:09 wm-bot2: Drained 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:05 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: fb99c967-b974-4314-a2fa-{{Gerrit|31ed0e883dd3}}, use this to unset). - cookbook ran by andrew@buster * 14:04 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 13:30 wm-bot2: Added 1 new OSDs ['cloudcephosd1031.eqiad.wmnet'] - cookbook ran by fran@wmf3169 * 13:30 wm-bot2: Added OSD cloudcephosd1031.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 13:26 wm-bot2: Adding OSD cloudcephosd1031.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 13:26 wm-bot2: Adding new OSDs ['cloudcephosd1031.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 13:16 wm-bot2: Adding OSD cloudcephosd1031.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 13:16 wm-bot2: Adding new OSDs ['cloudcephosd1031.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 13:08 wm-bot2: Adding OSD cloudcephosd1031.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 13:08 wm-bot2: Adding new OSDs ['cloudcephosd1031.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 12:45 wm-bot2: Finished rebooting node cloudcephosd1031.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 12:42 wm-bot2: Rebooting node cloudcephosd1031.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 12:41 wm-bot2: Adding OSD cloudcephosd1031.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 12:41 wm-bot2: Adding new OSDs ['cloudcephosd1031.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 12:26 arturo: [codfw1dev] cloudnet2005/2006-dev are now on a single NIC setup ([[phab:T318824|T318824]]) * 11:44 arturo: sysctl change (cleanup) on cloudnet1003/1004 === 2022-09-27 === * 10:48 wm-bot2: Added OSD cloudcephosd1030.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 10:35 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 10:32 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 10:32 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 10:32 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@wmf3169 * 10:26 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:23 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:22 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 10:22 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 10:17 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:14 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:14 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 10:14 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 10:09 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:05 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 10:05 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 10:05 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 10:05 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 10:05 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 === 2022-09-26 === * 18:37 wm-bot2: Safe reboot of 'cloudvirt1024.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 18:37 wm-bot2: Unset cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 18:33 wm-bot2: Drained 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:31 andrewbogott: rebooting cloudvirt1028 for [[phab:T317391|T317391]] * 18:28 wm-bot2: Safe reboot of 'cloudvirt-wdqs1002.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:28 wm-bot2: Unset cloudvirt 'cloudvirt-wdqs1002.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:28 wm-bot2: Safe reboot of 'cloudvirt-wdqs1003.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:28 wm-bot2: Unset cloudvirt 'cloudvirt-wdqs1003.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:28 wm-bot2: Safe reboot of 'cloudvirt-wdqs1001.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:28 wm-bot2: Unset cloudvirt 'cloudvirt-wdqs1001.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:25 wm-bot2: Drained 'cloudvirt-wdqs1003.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:25 wm-bot2: Set cloudvirt 'cloudvirt-wdqs1003.eqiad.wmnet' maintenance (downtime id: 6d1c4c53-76b9-493f-8c44-{{Gerrit|eb0413dfb9d0}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:25 wm-bot2: Drained 'cloudvirt-wdqs1002.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:25 wm-bot2: Set cloudvirt 'cloudvirt-wdqs1002.eqiad.wmnet' maintenance (downtime id: 6bb80b65-616c-4e45-be4f-{{Gerrit|d2cdd8abb2bf}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:25 wm-bot2: Drained 'cloudvirt-wdqs1001.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:25 wm-bot2: Set cloudvirt 'cloudvirt-wdqs1001.eqiad.wmnet' maintenance (downtime id: a3bba7e7-bcd3-482d-a486-{{Gerrit|06b6f53fd899}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:24 wm-bot2: Draining 'cloudvirt-wdqs1003.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:24 wm-bot2: Safe rebooting 'cloudvirt-wdqs1003.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:24 wm-bot2: Draining 'cloudvirt-wdqs1002.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:24 wm-bot2: Safe rebooting 'cloudvirt-wdqs1002.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:24 wm-bot2: Draining 'cloudvirt-wdqs1001.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:24 wm-bot2: Safe rebooting 'cloudvirt-wdqs1001.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:18 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 8fb59651-fd42-4aee-a978-{{Gerrit|36bc7f79b4d5}}, use this to unset). - cookbook ran by andrew@buster * 18:18 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:17 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:01 wm-bot2: Safe reboot of 'cloudvirt1023.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:01 wm-bot2: Unset cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:58 wm-bot2: Drained 'cloudvirt1023.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:51 wm-bot2: Safe reboot of 'cloudvirt1025.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:51 wm-bot2: Unset cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:47 wm-bot2: Safe reboot of 'cloudvirt1027.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 17:47 wm-bot2: Unset cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 17:47 wm-bot2: Drained 'cloudvirt1025.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:45 wm-bot2: Drained 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:38 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: abfc35bb-331e-4d7e-bb92-{{Gerrit|35e675abce32}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:37 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:37 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:37 wm-bot2: Safe reboot of 'cloudvirt1022.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:37 wm-bot2: Unset cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:34 wm-bot2: Drained 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:33 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: d4fe6ea8-76eb-45f7-b6cf-{{Gerrit|66f54ba9801c}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:33 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance (downtime id: c0671bda-9aae-4960-85be-{{Gerrit|58aaa9c8e021}}, use this to unset). - cookbook ran by andrew@buster * 17:32 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:32 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:32 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:32 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:31 wm-bot2: Safe reboot of 'cloudvirt1029.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 17:31 wm-bot2: Unset cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 17:31 wm-bot2: Safe reboot of 'cloudvirt1030.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:31 wm-bot2: Unset cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:28 wm-bot2: Drained 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:28 wm-bot2: Drained 'cloudvirt1030.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:24 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: be6559f9-4d72-4492-b9b9-{{Gerrit|23c636d6554e}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:23 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:23 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:21 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: 237dfb1b-5382-408d-aa35-{{Gerrit|1aa1cfe4c5af}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:20 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:20 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:20 wm-bot2: Safe reboot of 'cloudvirt1021.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:20 wm-bot2: Unset cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:17 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: 088b97cd-0cf2-4c27-a67f-{{Gerrit|077ffa1c1c5e}}, use this to unset). - cookbook ran by andrew@buster * 17:16 wm-bot2: Drained 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:16 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: 91625000-aa6d-4887-bf34-{{Gerrit|ef7785a4af4a}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:16 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:16 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:16 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:15 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:15 wm-bot2: Safe reboot of 'cloudvirt1017.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 17:15 wm-bot2: Unset cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 17:15 wm-bot2: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance (downtime id: 7364c4fe-9f6b-4539-b6fa-{{Gerrit|1e768b602a1b}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:14 wm-bot2: Draining 'cloudvirt1030.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:14 wm-bot2: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:12 wm-bot2: Drained 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:12 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: cb8f8a21-73c0-4f54-9412-{{Gerrit|074d82582cb4}}, use this to unset). - cookbook ran by andrew@buster * 17:11 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:11 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:09 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: d23dabd2-3bfc-4ce0-9533-{{Gerrit|cd1cf910f8e1}}, use this to unset). - cookbook ran by andrew@buster * 17:08 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:08 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:05 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: 4d7ddae1-f621-4dd5-8616-{{Gerrit|29b7699b692f}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:04 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:04 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:04 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: 4628dd6f-5b49-417f-b6ab-{{Gerrit|5c41992192b4}}, use this to unset). - cookbook ran by andrew@buster * 17:03 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:03 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:02 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: c95529cf-b2a5-4553-bd6e-{{Gerrit|6ad4dcb2a105}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:01 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: 36c68112-a964-4a7c-a093-{{Gerrit|8a6d481dea7d}}, use this to unset). - cookbook ran by andrew@buster * 17:01 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:01 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 17:00 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 17:00 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 16:49 wm-bot2: Safe reboot of 'cloudvirt1050.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 16:49 wm-bot2: Unset cloudvirt 'cloudvirt1050.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 16:45 wm-bot2: Drained 'cloudvirt1050.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 16:31 wm-bot2: Set cloudvirt 'cloudvirt1050.eqiad.wmnet' maintenance (downtime id: 1244c159-65bb-476e-a702-{{Gerrit|8f43c253e499}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 16:30 wm-bot2: Draining 'cloudvirt1050.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 16:30 wm-bot2: Safe rebooting 'cloudvirt1050.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:22 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 13:19 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@wmf3169 * 13:18 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@wmf3169 * 13:18 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@wmf3169 * 12:32 dcaro: Changed the collation of labsdbaccount db to utf8mb4_bin ([[phab:T318047|T318047]]) * 10:14 arturo: deployed new version of maintain-dbusers ([[phab:T318047|T318047]]) === 2022-09-25 === * 15:06 wm-bot2: Drained 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:06 wm-bot2: Set cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance (downtime id: a050e47f-3a63-41ff-accb-{{Gerrit|3993c1e8b593}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:05 wm-bot2: Draining 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:05 wm-bot2: Safe rebooting 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:03 wm-bot2: Safe reboot of 'cloudvirt1052.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:03 wm-bot2: Unset cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:58 wm-bot2: Drained 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:58 wm-bot2: Set cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance (downtime id: 1a49d773-4dc9-4a6f-bd49-{{Gerrit|8663db77ce66}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:57 wm-bot2: Draining 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:57 wm-bot2: Safe rebooting 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:54 wm-bot2: Set cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance (downtime id: ab52dd42-68a6-4159-b345-{{Gerrit|e5625d209b57}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:53 wm-bot2: Draining 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:53 wm-bot2: Safe rebooting 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:37 wm-bot2: Set cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance (downtime id: 10337415-3095-4834-8538-{{Gerrit|b3a64bd6b18d}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:37 wm-bot2: Draining 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:37 wm-bot2: Safe rebooting 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:53 wm-bot2: Drained 'cloudvirt1053.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:53 wm-bot2: Set cloudvirt 'cloudvirt1053.eqiad.wmnet' maintenance (downtime id: 4cdbed3a-3775-4913-8c3b-{{Gerrit|afd57ae1ca55}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:52 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:52 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster === 2022-09-24 === * 17:37 andrewbogott: restarting neutron api on cloudcontrol1006; cause of outage unknown * 17:35 andrewbogott: restarting neutron-linuxbridge-agent on cloudvirt1022 === 2022-09-22 === * 15:14 wm-bot2: Safe reboot of 'cloudvirt1025.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:14 wm-bot2: Unset cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:10 wm-bot2: Drained 'cloudvirt1025.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:10 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 04849437-a304-45a1-a037-{{Gerrit|538741a2f801}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:09 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:09 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:07 wm-bot2: Safe reboot of 'cloudvirt1017.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:07 wm-bot2: Unset cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:04 wm-bot2: Drained 'cloudvirt1017.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:03 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: 8810a919-61e7-4ba5-af7f-{{Gerrit|c41d1e1c8c8b}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:03 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 15:03 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:59 wm-bot2: Safe reboot of 'cloudvirt1021.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:59 wm-bot2: Unset cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:56 wm-bot2: Drained 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:42 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: b6f8e0b6-f35c-41fd-8035-{{Gerrit|6efde61db3a3}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:42 wm-bot2: Safe reboot of 'cloudvirt1027.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:42 wm-bot2: Unset cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:42 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:42 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:41 wm-bot2: Safe reboot of 'cloudvirt1022.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:41 wm-bot2: Unset cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:41 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: 784bff2c-8160-44dc-9bcf-{{Gerrit|ff23055a0527}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:40 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:40 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:39 wm-bot2: Drained 'cloudvirt1027.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:38 wm-bot2: Drained 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:38 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: 21152e9e-bc67-4024-b68e-{{Gerrit|fd671c398642}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:37 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:37 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:37 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance (downtime id: f07ddb1d-e8da-43f3-8140-{{Gerrit|bc66789d37c8}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:36 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:36 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:33 wm-bot2: Safe reboot of 'cloudvirt1024.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:33 wm-bot2: Unset cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:30 wm-bot2: Drained 'cloudvirt1024.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:16 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 4e4b03fc-e3f1-440e-9097-{{Gerrit|cce5edaa1f08}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:16 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance (downtime id: d4bb2160-f492-4ddc-a5b7-{{Gerrit|eeee37007d14}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:16 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: 9bf00a6a-3697-4201-9a6e-{{Gerrit|76d499eccfa1}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:15 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:13 wm-bot2: Safe reboot of 'cloudvirt1030.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 14:13 wm-bot2: Unset cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:47 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: f2888490-1804-4b23-b7b0-{{Gerrit|677853fb9ef7}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:46 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:46 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:46 wm-bot2: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance (downtime id: 8a3e8b0a-ced3-4667-a121-{{Gerrit|66df673c7ed8}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:46 wm-bot2: Safe reboot of 'cloudvirt1033.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:46 wm-bot2: Unset cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:45 wm-bot2: Draining 'cloudvirt1030.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:45 wm-bot2: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:45 wm-bot2: Safe reboot of 'cloudvirt1032.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:45 wm-bot2: Unset cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:43 wm-bot2: Set cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance (downtime id: 35c761b3-f0de-4b23-9b11-{{Gerrit|2ed2e474c89f}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:42 wm-bot2: Draining 'cloudvirt1031.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:42 wm-bot2: Safe rebooting 'cloudvirt1031.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:42 wm-bot2: Drained 'cloudvirt1033.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:42 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: 6121a0e1-754a-47fa-b51b-{{Gerrit|265f6386f396}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:41 wm-bot2: Drained 'cloudvirt1032.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:41 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:41 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:40 wm-bot2: Safe reboot of 'cloudvirt1035.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:40 wm-bot2: Unset cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:39 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: 189ae1bb-c69d-4eb0-aee9-{{Gerrit|90ab1f492aa8}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:38 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:38 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:36 wm-bot2: Drained 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:35 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: acc53776-68bf-47a4-bb82-{{Gerrit|c571b3df2945}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:35 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:35 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:31 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 07770ae7-3e7e-4615-8174-{{Gerrit|513767f95b1d}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:30 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:30 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:29 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 7ae20a7d-8bff-4bc2-91c5-{{Gerrit|b0005c8164d1}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:29 wm-bot2: Set cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance (downtime id: b596f855-8019-49d2-9657-{{Gerrit|0d513ad7acb5}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:29 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:29 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:28 wm-bot2: Draining 'cloudvirt1032.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:28 wm-bot2: Safe rebooting 'cloudvirt1032.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:28 wm-bot2: Safe reboot of 'cloudvirt1034.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:28 wm-bot2: Unset cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:25 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: 9172c785-d2d6-46e5-bed3-{{Gerrit|2f49d1033f78}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:24 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:24 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:24 wm-bot2: Drained 'cloudvirt1034.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:23 wm-bot2: Safe reboot of 'cloudvirt1036.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:23 wm-bot2: Unset cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:19 wm-bot2: Drained 'cloudvirt1036.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:04 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance (downtime id: 85b4e0e2-e635-485f-9ce5-{{Gerrit|341981b3887f}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:04 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 9494c283-02c8-42da-98c3-{{Gerrit|139b4d119ef8}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:04 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: d181f301-249f-4b5a-bd85-{{Gerrit|47d87262d0ed}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:03 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:03 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:03 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:03 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:03 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 13:03 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster === 2022-09-20 === * 21:02 wm-bot2: Safe reboot of 'cloudvirt1037.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 21:02 wm-bot2: Unset cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:58 wm-bot2: Drained 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:58 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: 1c4246a4-9cee-4423-8cd1-{{Gerrit|4f52f61d503f}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:57 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:57 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:57 wm-bot2: Safe reboot of 'cloudvirt1038.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:57 wm-bot2: Unset cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:53 wm-bot2: Drained 'cloudvirt1038.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:50 wm-bot2: Safe reboot of 'cloudvirt1039.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:50 wm-bot2: Unset cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:49 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: 82665ecc-4431-48fe-b255-{{Gerrit|7e9d518be217}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:48 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:48 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:47 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: cd284562-e3b0-4504-9977-{{Gerrit|2b18032bbf3c}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:46 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:46 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:46 wm-bot2: Drained 'cloudvirt1039.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:45 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: 8fb2a646-16a5-4183-94a1-{{Gerrit|ee6bbbf029fa}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:44 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:44 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:32 wm-bot2: Set cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance (downtime id: d561fe45-1582-4cd8-bbc9-{{Gerrit|a356c39f3330}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:32 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance (downtime id: 351c365c-0228-4907-a279-{{Gerrit|01795b1e62ce}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:32 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: 3beec1dd-0132-4f5c-adce-{{Gerrit|5a7dc56060b6}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:31 wm-bot2: Draining 'cloudvirt1039.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:31 wm-bot2: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:31 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:31 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:31 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:31 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:30 wm-bot2: Safe reboot of 'cloudvirt1040.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:29 wm-bot2: Unset cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:28 wm-bot2: Safe reboot of 'cloudvirt1041.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:28 wm-bot2: Unset cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:26 wm-bot2: Drained 'cloudvirt1040.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:24 wm-bot2: Drained 'cloudvirt1041.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:20 wm-bot2: Safe reboot of 'cloudvirt1042.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:20 wm-bot2: Unset cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:16 wm-bot2: Drained 'cloudvirt1042.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:07 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance (downtime id: 353a8527-ad5e-4898-937c-{{Gerrit|303d16801e28}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:07 wm-bot2: Set cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance (downtime id: fc235146-f070-4723-9503-{{Gerrit|e20cbb877755}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:07 wm-bot2: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance (downtime id: cdf712ac-c083-4a63-af83-{{Gerrit|514ca5cdc76d}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:06 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:06 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:06 wm-bot2: Draining 'cloudvirt1041.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:06 wm-bot2: Safe rebooting 'cloudvirt1041.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:06 wm-bot2: Draining 'cloudvirt1040.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:06 wm-bot2: Safe rebooting 'cloudvirt1040.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:05 wm-bot2: Safe reboot of 'cloudvirt1043.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:05 wm-bot2: Unset cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:03 wm-bot2: Safe reboot of 'cloudvirt1044.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:03 wm-bot2: Unset cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:02 wm-bot2: Safe reboot of 'cloudvirt1045.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:02 wm-bot2: Unset cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:01 wm-bot2: Drained 'cloudvirt1043.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:01 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance (downtime id: 8f2d2d32-e99e-4e2c-9472-{{Gerrit|ca33b577d01f}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:00 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:00 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:00 wm-bot2: Drained 'cloudvirt1044.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:00 wm-bot2: Set cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance (downtime id: 66ee7d87-cee1-4c8f-b5c3-{{Gerrit|5f58a6cea336}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:59 wm-bot2: Draining 'cloudvirt1044.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:59 wm-bot2: Safe rebooting 'cloudvirt1044.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:58 wm-bot2: Drained 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:58 wm-bot2: Drained 'cloudvirt1045.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:58 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance (downtime id: 69035b38-9910-46f1-9582-{{Gerrit|3105c1ba4d16}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:57 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:57 wm-bot2: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:56 wm-bot2: Safe reboot of 'cloudvirt1046.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:56 wm-bot2: Unset cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:52 wm-bot2: Drained 'cloudvirt1046.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:52 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: 5a94c9d1-c956-4b03-a92b-{{Gerrit|b112810906a4}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:51 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:51 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:50 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: 216a9120-4b9a-4be3-99db-{{Gerrit|3fd5fdc25146}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:49 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:49 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:48 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: 46fb1964-3cf4-47fd-ad79-{{Gerrit|475f33c8ed38}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:48 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:48 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:47 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: b04a4f8c-db82-4258-8608-{{Gerrit|2a78ec9193ed}}, use this to unset). - cookbook ran by andrew@buster * 19:46 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:46 wm-bot2: Drained 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:46 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance (downtime id: 7ae7c258-ecb8-47d1-a991-{{Gerrit|3fa617e305bf}}, use this to unset). - cookbook ran by andrew@buster * 19:45 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:44 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance (downtime id: 377fb0e5-132e-4e7a-9079-{{Gerrit|37937c3d42bf}}, use this to unset). - cookbook ran by andrew@buster * 19:44 wm-bot2: Set cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance (downtime id: 6c4222ae-796e-4890-8820-{{Gerrit|9e7ce5438f2a}}, use this to unset). - cookbook ran by andrew@buster * 19:43 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:43 wm-bot2: Draining 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:38 wm-bot2: Drained 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:34 wm-bot2: Safe reboot of 'cloudvirt1047.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:34 wm-bot2: Unset cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:31 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: eff98e2e-ae99-41d0-a68e-{{Gerrit|671d941fcf21}}, use this to unset). - cookbook ran by andrew@buster * 19:30 wm-bot2: Drained 'cloudvirt1047.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:30 wm-bot2: Set cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance (downtime id: 5310f1eb-0425-4f59-8fb5-{{Gerrit|1d9c6f6d2f23}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:30 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:30 wm-bot2: Draining 'cloudvirt1047.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:30 wm-bot2: Safe rebooting 'cloudvirt1047.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:29 wm-bot2: Drained 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:25 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance (downtime id: d8ab4682-026a-4ddc-bfc5-{{Gerrit|48166bd9789a}}, use this to unset). - cookbook ran by andrew@buster * 19:24 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:23 wm-bot2: Safe reboot of 'cloudvirt1048.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:23 wm-bot2: Unset cloudvirt 'cloudvirt1048.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:19 wm-bot2: Drained 'cloudvirt1048.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:19 wm-bot2: Set cloudvirt 'cloudvirt1048.eqiad.wmnet' maintenance (downtime id: eb2f9d94-8ea5-48d5-a336-{{Gerrit|97dd407b1ce3}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:19 wm-bot2: Draining 'cloudvirt1048.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:19 wm-bot2: Safe rebooting 'cloudvirt1048.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:18 wm-bot2: Drained 'cloudvirt1048.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:14 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: 305fe2d7-913e-4638-95c2-{{Gerrit|4e810dc6b2a3}}, use this to unset). - cookbook ran by andrew@buster * 19:14 wm-bot2: Set cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance (downtime id: 5bb3a66b-9450-4230-8d09-{{Gerrit|d78473d72ba2}}, use this to unset). - cookbook ran by andrew@buster * 19:13 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:13 wm-bot2: Draining 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:07 wm-bot2: Set cloudvirt 'cloudvirt1048.eqiad.wmnet' maintenance (downtime id: b1e39b49-3a88-4e42-913c-{{Gerrit|9892eafad447}}, use this to unset). - cookbook ran by andrew@buster * 19:06 wm-bot2: Draining 'cloudvirt1048.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:05 andrewbogott: putting cloudvirt1049-1052 into 'ceph' pool, taking out of 'spare' pool. cloudvirt1053 will remain our only spare. * 19:02 wm-bot2: Safe reboot of 'cloudvirt1049.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:02 wm-bot2: Unset cloudvirt 'cloudvirt1049.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:00 wm-bot2: Safe reboot of 'cloudvirt1050.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:00 wm-bot2: Unset cloudvirt 'cloudvirt1050.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:59 wm-bot2: Safe reboot of 'cloudvirt1051.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:59 wm-bot2: Unset cloudvirt 'cloudvirt1051.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:58 wm-bot2: Drained 'cloudvirt1049.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:58 wm-bot2: Set cloudvirt 'cloudvirt1049.eqiad.wmnet' maintenance (downtime id: 8f24ea75-e107-4a13-8bd6-{{Gerrit|950b941b8e03}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:58 wm-bot2: Draining 'cloudvirt1049.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:58 wm-bot2: Safe rebooting 'cloudvirt1049.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:57 wm-bot2: Safe reboot of 'cloudvirt1052.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:57 wm-bot2: Unset cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:56 wm-bot2: Drained 'cloudvirt1050.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:56 wm-bot2: Set cloudvirt 'cloudvirt1050.eqiad.wmnet' maintenance (downtime id: 5c3e8fa4-cd1e-43f4-a423-{{Gerrit|ba7fe2459d4e}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:55 wm-bot2: Drained 'cloudvirt1051.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:55 wm-bot2: Set cloudvirt 'cloudvirt1051.eqiad.wmnet' maintenance (downtime id: e3fe22d6-932c-4502-8838-{{Gerrit|82cdc4c8f1c6}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:55 wm-bot2: Draining 'cloudvirt1050.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:55 wm-bot2: Safe rebooting 'cloudvirt1050.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:54 wm-bot2: Draining 'cloudvirt1051.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:54 wm-bot2: Safe rebooting 'cloudvirt1051.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:53 wm-bot2: Safe reboot of 'cloudvirt1053.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:53 wm-bot2: Unset cloudvirt 'cloudvirt1053.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:52 wm-bot2: Drained 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:52 wm-bot2: Set cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance (downtime id: fd824b92-b6ff-4631-a2d4-{{Gerrit|debb25d48223}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:52 wm-bot2: Draining 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:52 wm-bot2: Safe rebooting 'cloudvirt1052.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:51 wm-bot2: Drained 'cloudvirt1052.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:49 wm-bot2: Set cloudvirt 'cloudvirt1052.eqiad.wmnet' maintenance (downtime id: 52d5cd5b-0f56-453a-85f8-{{Gerrit|0b5e656ae7f9}}, use this to unset). - cookbook ran by andrew@buster * 18:49 wm-bot2: Drained 'cloudvirt1053.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:49 wm-bot2: Set cloudvirt 'cloudvirt1053.eqiad.wmnet' maintenance (downtime id: b92f66a3-a487-4515-b68b-{{Gerrit|1c257a57b893}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:48 wm-bot2: Draining 'cloudvirt1052.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:48 wm-bot2: Draining 'cloudvirt1053.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 18:48 wm-bot2: Safe rebooting 'cloudvirt1053.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster === 2022-09-19 === * 20:07 wm-bot2: Safe reboot of 'cloudvirt1026.eqiad.wmnet' finished successfully. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:07 wm-bot2: Unset cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:03 wm-bot2: Drained 'cloudvirt1026.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:03 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance (downtime id: 4bfca6cd-dfbc-4a79-9203-{{Gerrit|b432bfeee5c2}}, use this to unset). ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:02 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 20:02 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. ([[phab:T317391|T317391]]) - cookbook ran by andrew@buster * 19:41 wm-bot2: Drained 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:21 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance (downtime id: 70479325-609b-4349-9094-{{Gerrit|739eb9b3bb30}}, use this to unset). - cookbook ran by andrew@buster * 19:21 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster === 2022-09-14 === * 16:57 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@Francesco’s-MacBook-Pro * 16:53 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@Francesco’s-MacBook-Pro * 16:53 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@Francesco’s-MacBook-Pro * 16:53 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@Francesco’s-MacBook-Pro * 11:01 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 10:58 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 10:57 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 10:57 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 10:09 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 10:09 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 10:06 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 10:06 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 09:46 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:43 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:43 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:43 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@Francesco’s-MacBook-Pro === 2022-09-13 === * 12:16 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 12:16 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 12:11 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 12:11 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 12:10 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by dcaro@vulcanus * 12:10 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by dcaro@vulcanus * 10:40 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:37 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:37 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:37 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:36 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster - cookbook ran by fran@Francesco’s-MacBook-Pro === 2022-09-10 === * 15:37 andrewbogott: restarting nova-conductor service (and possibly others, in response to lots of unanswered rabbitmq messages) === 2022-09-08 === * 19:12 andrewbogott: restarting nginx on proxy-03.project-proxy.eqiad1.wikimedia.cloud === 2022-09-07 === * 10:18 wm-bot2: Added OSD cloudcephosd1032.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:14 wm-bot2: Finished rebooting node cloudcephosd1032.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:10 wm-bot2: Rebooting node cloudcephosd1032.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:10 wm-bot2: Adding OSD cloudcephosd1032.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:10 wm-bot2: Adding new OSDs ['cloudcephosd1032.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:41 dhinus: Temporarily removing cloudcephosd1030 from Ceph cluster (https://phabricator.wikimedia.org/T314870) === 2022-08-30 === * 14:59 andrewbogott: manually marking most eqiad1 cloud* servers down in icinga for [[phab:T296561|T296561]] * 10:43 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:39 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:38 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:38 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:39 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:32 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:32 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:32 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:14 wm-bot2: Finished rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:11 wm-bot2: Rebooting node cloudcephosd1030.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:10 wm-bot2: Adding OSD cloudcephosd1030.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:10 wm-bot2: Adding new OSDs ['cloudcephosd1030.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro === 2022-08-25 === * 15:14 wm-bot2: Added 1 new OSDs ['cloudcephosd1029.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 15:14 wm-bot2: Added OSD cloudcephosd1029.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 15:02 wm-bot2: Finished rebooting node cloudcephosd1029.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 14:59 wm-bot2: Rebooting node cloudcephosd1029.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 14:58 wm-bot2: Adding OSD cloudcephosd1029.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 14:58 wm-bot2: Adding new OSDs ['cloudcephosd1029.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro === 2022-08-24 === * 22:07 andrewbogott: replaced cloudservices1003 with cloudservices1005 [[phab:T304888|T304888]] * 10:45 wm-bot2: Added 1 new OSDs ['cloudcephosd1028.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:45 wm-bot2: Added OSD cloudcephosd1028.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:37 wm-bot2: Finished rebooting node cloudcephosd1028.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:34 wm-bot2: Rebooting node cloudcephosd1028.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:33 wm-bot2: Adding OSD cloudcephosd1028.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 10:33 wm-bot2: Adding new OSDs ['cloudcephosd1028.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro === 2022-08-23 === * 13:46 wm-bot2: Added 1 new OSDs ['cloudcephosd1027.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:46 wm-bot2: Added OSD cloudcephosd1027.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:27 wm-bot2: Finished rebooting node cloudcephosd1027.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:24 wm-bot2: Rebooting node cloudcephosd1027.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:22 wm-bot2: Adding OSD cloudcephosd1027.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:22 wm-bot2: Adding new OSDs ['cloudcephosd1027.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro === 2022-08-21 === * 21:12 andrewbogott: restarted neutron-dhcp-agent on cloudnet1003. it was claiming to be unable to contact Rabbit but seems happy after a restart === 2022-08-20 === * 07:39 dcaro_away: cloudvirt1023 is back up, VMs are starting to recover ([[phab:T315718|T315718]]) * 07:23 dcaro_away: cloudvirt1023 seems to have gotten some hardware issue from racadm lclog view "System CPU Resetting.", rebooting and doing memory checks ([[phab:T315718|T315718]]) === 2022-08-19 === * 17:06 taavi: [codfw1dev] restart mariadb on clouddb2002-dev to pick up certificate config changes [[phab:T310795|T310795]] === 2022-08-18 === * 13:25 wm-bot2: Added 1 new OSDs ['cloudcephosd1026.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:25 wm-bot2: Added OSD cloudcephosd1026.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:15 wm-bot2: Finished rebooting node cloudcephosd1026.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:12 wm-bot2: Rebooting node cloudcephosd1026.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:10 wm-bot2: Adding OSD cloudcephosd1026.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 13:10 wm-bot2: Adding new OSDs ['cloudcephosd1026.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 07:29 dcaro: Starting up all the osd daemons on cloudcephosd1025 ([[phab:T314870|T314870]]) === 2022-08-17 === * 10:50 wm-bot2: Added 1 new OSDs ['cloudcephosd1025.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 10:49 wm-bot2: Added OSD cloudcephosd1025.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 10:40 wm-bot2: Finished rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 10:37 wm-bot2: Rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 10:37 wm-bot2: Adding OSD cloudcephosd1025.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 10:37 wm-bot2: Adding new OSDs ['cloudcephosd1025.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 09:50 wm-bot2: Rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 09:49 wm-bot2: Adding OSD cloudcephosd1025.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 09:49 wm-bot2: Adding new OSDs ['cloudcephosd1025.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@foz * 09:16 wm-bot2: Rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:16 wm-bot2: Adding OSD cloudcephosd1025.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro * 09:16 wm-bot2: Adding new OSDs ['cloudcephosd1025.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@Francesco’s-MacBook-Pro === 2022-08-16 === * 22:39 andrewbogott: replacing the now-rebuilt cloudvirt1025 in 'ceph' aggregate and removing it from the 'maintenance' aggregate * 17:41 andrewbogott: removing cloudvirt1025 from the 'ceph' aggregate and adding it to the 'maintenance' aggregate * 17:40 andrewbogott: reimaging cloudvirt1025 after I accidentally deleted the hw raid * 17:38 andrewbogott: root@cloudcontrol1005:~# cinder-manage volume update_host --currenthost cloudcontrol1003@rbd#RBD --newhost cloudcontrol1005@rbd#RBD * 17:37 andrewbogott: root@cloudcontrol1005:~# cinder-manage volume update_host --currenthost cloudcontrol1004@rbd#RBD --newhost cloudcontrol1006@rbd#RBD * 16:26 wm-bot2: Ceph cluster at eqiad1 set out of maintenance. - cookbook ran by dcaro@vulcanus * 15:43 wm-bot2: Restarting the osd daemons from nodes cloudcephosd1001,cloudcephosd1002,cloudcephosd1003,cloudcephosd1004,cloudcephosd1005,cloudcephosd1006,cloudcephosd1007,cloudcephosd1008,cloudcephosd1009,cloudcephosd1010,cloudcephosd1011,cloudcephosd1012,cloudcephosd1013,cloudcephosd1014,cloudcephosd1015,cloudcephosd1016,cloudcephosd1017,cloudcephosd1018,cloudcephosd1019,cloudcephosd1020,cloudcephosd1021,cloudcephosd1022,cloudcephosd1023,c * 15:42 wm-bot2: Finished restarting all the OSD daemons from the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] - cookbook ran by dcaro@vulcanus * 15:38 wm-bot2: Restarting the osd daemons from nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev - cookbook ran by dcaro@vulcanus * 13:08 wm-bot2: Restarting the osd daemons from nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev - cookbook ran by dcaro@vulcanus * 13:07 wm-bot2: Restarting the osd daemons from nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev - cookbook ran by dcaro@vulcanus * 13:02 wm-bot2: Restarting the osd daemons from nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev - cookbook ran by dcaro@vulcanus * 13:01 wm-bot2: Restarting the osd daemons from nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev - cookbook ran by dcaro@vulcanus * 12:59 wm-bot2: Restarting the osd daemons from nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev - cookbook ran by dcaro@vulcanus === 2022-08-14 === * 18:36 taavi: deleted the http keystone endpoints from the keystone service catalog === 2022-08-11 === * 13:57 andrewbogott: decommissioning cloudcontrol1003 + cloudcontrl1004. I backed up $home in case anyone needs their files. * 08:42 wm-bot2: The cluster is now rebalanced after adding the new OSDs ['cloudcephosd1025.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 08:42 wm-bot2: Added 1 new OSDs ['cloudcephosd1025.eqiad.wmnet'] ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 08:42 wm-bot2: Added OSD cloudcephosd1025.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 08:40 wm-bot2: Finished rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 08:36 wm-bot2: Rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 08:36 wm-bot2: Adding OSD cloudcephosd1025.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 08:36 wm-bot2: Adding new OSDs ['cloudcephosd1025.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station === 2022-08-10 === * 13:10 wm-bot2: Finished rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 13:06 wm-bot2: Rebooting node cloudcephosd1025.eqiad.wmnet ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 13:06 wm-bot2: Adding OSD cloudcephosd1025.eqiad.wmnet... (1/1) ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station * 13:06 wm-bot2: Adding new OSDs ['cloudcephosd1025.eqiad.wmnet'] to the cluster ([[phab:T314870|T314870]]) - cookbook ran by fran@MacBook-Pro.station === 2022-08-04 === * 17:16 taavi: deleted all scheduler_fanout_ rabbit queues in an attempt to fix scheduling * 16:32 taavi: restart neutron-l3-agent to pick up rabbit config changes * 15:12 andrewbogott: stopping rabbitmq on cloudcontrol1xxx * 09:57 taavi: stop wikitech_run_jobs.timer on labweb1001/1002, hosts pending decom === 2022-08-03 === * 20:55 andrewbogott: root@tools-checker-04:~# systemctl restart uwsgi-toolschecker_cron.service * 20:41 andrewbogott: restarting neutron-l3-agent.service on cloudnet1003 and 1004. The agent was routing properly but had lost touch with rabbitmq === 2022-08-02 === * 14:07 andrewbogott: shutting down codfw1dev ceph cluster according to https://docs.mirantis.com/mcp/q4-18/mcp-operations-guide/scheduled-maintenance-power-outage/power-off-ceph-cluster.html * 13:54 andrewbogott: shutting down basically all of codfw1dev to support pdu maintenance -- all the ceph OSDs will lose power so best to have everything stopped. === 2022-07-27 === * 19:32 andrewbogott: switching the openstack.eqiad1.wikimedia.cloud endpoint from cloudcontrol1004 to 1006, https://gerrit.wikimedia.org/r/c/operations/dns/+/817878/2/templates/wikimediacloud.org#54 * 16:33 andrewbogott: here is a test message in the admin channel === 2022-07-25 === * 13:43 andrewbogott: pooling cloudweb100[34] and depooling labweb100[12] for testing in prep for decomming labweb100[12] === 2022-07-22 === * 16:41 taavi: depool cloudweb1003/1004 since horizon seems to be having issues * 16:22 taavi: pooling cloudweb1003/1004 now that grant issues are sorted === 2022-07-21 === * 18:26 andrewbogott: depooling cloudweb1003 and 1004 for wikitech, horizon, striker -- pending db grant changes * 18:06 andrewbogott: pooling cloudweb1003 and 1004 for wikitech, horizon, striker === 2022-07-20 === * 18:02 dcaro: things seem stable, trying to bring up a the last rabbit node, cloudcontrol1007 ([[phab:T313400|T313400]]) * 17:45 bd808: `sudo service striker restart` on labweb1002 * 17:43 bd808: `sudo service striker restart` on labweb1001 * 17:10 dcaro: things seem stable, trying to bring up a fourth rabbit node, cloudcontrol1006 ([[phab:T313400|T313400]]) * 16:26 dcaro: things seem stable, trying to bring up a third, cloudcontrol1005 ([[phab:T313400|T313400]]) * 15:51 dcaro: things seem stable now with one rabbit node, trying to bring up a second ([[phab:T313400|T313400]]) * 14:16 dcaro: stopping rabbin on cloudcontrol1004, leaving only 1003 alive ([[phab:T313400|T313400]]) * 13:17 dcaro: restarting the whole rabbit cluster ([[phab:T313400|T313400]]) === 2022-07-19 === * 16:30 wm-bot2: Safe reboot of 'cloudvirt1045.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 16:30 wm-bot2: Unset cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 16:26 wm-bot2: Drained 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 16:18 wm-bot2: Safe reboot of 'cloudvirt1044.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 16:18 wm-bot2: Unset cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 16:14 wm-bot2: Drained 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 16:01 wm-bot2: Safe reboot of 'cloudvirt1047.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 16:01 wm-bot2: Unset cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 15:57 wm-bot2: Drained 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:57 wm-bot2: Set cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance (downtime id: 3da2d4f6-5c5b-4a21-9a0b-{{Gerrit|2b010960ed6a}}, use this to unset). - cookbook ran by andrew@buster * 15:56 wm-bot2: Draining 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:56 wm-bot2: Safe rebooting 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:56 wm-bot2: Safe reboot of 'cloudvirt1046.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 15:56 wm-bot2: Unset cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 15:52 wm-bot2: Drained 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:49 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance (downtime id: 8a10505a-107d-4e78-ba84-{{Gerrit|363a7dea8d69}}, use this to unset). - cookbook ran by andrew@buster * 15:47 wm-bot2: Set cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance (downtime id: 752aa58b-44d4-4340-b05d-{{Gerrit|911b14f2314d}}, use this to unset). - cookbook ran by andrew@buster * 15:47 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance (downtime id: f89f0851-ce3f-4a92-899e-{{Gerrit|0d4638acc9c3}}, use this to unset). - cookbook ran by andrew@buster * 15:46 wm-bot2: Draining 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:46 wm-bot2: Safe rebooting 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:46 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:46 wm-bot2: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:46 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:46 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:45 andrewbogott: adding new hosts to the 'ceph' aggregate: cloudvirt1046, 1047, 1048 * 15:44 wm-bot2: Safe reboot of 'cloudvirt1041.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 15:44 wm-bot2: Unset cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 15:42 wm-bot2: Safe reboot of 'cloudvirt1043.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 15:42 wm-bot2: Unset cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 15:40 wm-bot2: Drained 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:38 wm-bot2: Drained 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:37 wm-bot2: Safe reboot of 'cloudvirt1042.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 15:37 wm-bot2: Unset cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 15:33 wm-bot2: Drained 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:32 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance (downtime id: 32d9fdb7-6c42-4cf4-9950-{{Gerrit|6e2442255161}}, use this to unset). - cookbook ran by andrew@buster * 15:31 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:31 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:16 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance (downtime id: 56ce1388-792c-442c-bb89-{{Gerrit|e4c3869b0acf}}, use this to unset). - cookbook ran by andrew@buster * 15:15 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:15 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:14 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance (downtime id: 3f84bf2e-24de-4ebe-8117-{{Gerrit|0f4a5de29d09}}, use this to unset). - cookbook ran by andrew@buster * 15:14 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:14 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:13 wm-bot2: Safe reboot of 'cloudvirt1040.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 15:13 wm-bot2: Unset cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 15:12 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance (downtime id: 5418e923-997c-49ae-b937-{{Gerrit|35fc0c3038eb}}, use this to unset). - cookbook ran by andrew@buster * 15:12 wm-bot2: Set cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance (downtime id: c3dfdc2b-0725-4d01-87f2-{{Gerrit|440a5ac4694c}}, use this to unset). - cookbook ran by andrew@buster * 15:11 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:11 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:11 wm-bot2: Draining 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:11 wm-bot2: Safe rebooting 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:09 wm-bot2: Drained 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:04 wm-bot2: Safe reboot of 'cloudvirt1039.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 15:04 wm-bot2: Unset cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 15:00 wm-bot2: Drained 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:54 wm-bot2: Safe reboot of 'cloudvirt1038.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 14:54 wm-bot2: Unset cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 14:50 wm-bot2: Drained 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:46 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance (downtime id: 49dbf7de-c58a-4b68-bc7b-{{Gerrit|f001a047f4c7}}, use this to unset). - cookbook ran by andrew@buster * 14:46 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:46 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:44 wm-bot2: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance (downtime id: a4e95168-39b5-452d-8fce-{{Gerrit|7875ae44a62b}}, use this to unset). - cookbook ran by andrew@buster * 14:44 wm-bot2: Draining 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:44 wm-bot2: Safe rebooting 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:43 wm-bot2: Set cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance (downtime id: b46234ee-9900-4708-a815-{{Gerrit|4324379be48c}}, use this to unset). - cookbook ran by andrew@buster * 14:43 wm-bot2: Safe reboot of 'cloudvirt1036.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 14:43 wm-bot2: Unset cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 14:42 wm-bot2: Draining 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:42 wm-bot2: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:39 wm-bot2: Safe reboot of 'cloudvirt1037.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 14:39 wm-bot2: Unset cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 14:38 wm-bot2: Drained 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:38 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance (downtime id: 12bc87a9-7602-4795-818c-{{Gerrit|f7151b1c9ead}}, use this to unset). - cookbook ran by andrew@buster * 14:37 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:37 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:35 wm-bot2: Drained 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:32 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance (downtime id: ba787a26-62db-4662-ade2-{{Gerrit|2e3061b7390d}}, use this to unset). - cookbook ran by andrew@buster * 14:31 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:31 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:29 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: 6f0e79c6-2d91-478e-92df-{{Gerrit|0f65de48212f}}, use this to unset). - cookbook ran by andrew@buster * 14:29 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance (downtime id: 7f577c0f-0182-498f-b198-{{Gerrit|307d51df8c3a}}, use this to unset). - cookbook ran by andrew@buster * 14:28 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:28 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:28 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:28 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:28 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance (downtime id: bd1bff92-bfb8-483d-a719-{{Gerrit|13fbdf3d8b21}}, use this to unset). - cookbook ran by andrew@buster * 14:27 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:27 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:18 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance (downtime id: c9cd7dbb-c5ad-4364-a8e0-{{Gerrit|afb52ef98683}}, use this to unset). - cookbook ran by andrew@buster * 14:18 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance (downtime id: 4deca8a4-f73f-4b42-b4be-{{Gerrit|acfe61421ed0}}, use this to unset). - cookbook ran by andrew@buster * 14:18 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance (downtime id: 6d0d58b2-8450-480d-94f2-{{Gerrit|93a257f25575}}, use this to unset). - cookbook ran by andrew@buster * 14:17 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:17 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:17 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:17 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:17 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:17 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 14:13 dcaro: deleting all the leftover fullstack images (was due to max number of mysql connections reached) * 04:51 wm-bot2: Safe reboot of 'cloudvirt1035.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 04:51 wm-bot2: Unset cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:47 wm-bot2: Drained 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:46 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 432747a2-9a15-44d4-ba8a-{{Gerrit|4e7e87eb8b76}}, use this to unset). - cookbook ran by andrew@buster * 04:45 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:45 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:45 wm-bot2: Safe reboot of 'cloudvirt1033.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 04:45 wm-bot2: Unset cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:45 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 1c5adb8c-efcf-4036-bf22-{{Gerrit|0eef364ae399}}, use this to unset). - cookbook ran by andrew@buster * 04:44 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:44 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:41 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 26feb897-9eb0-40b6-9525-{{Gerrit|4cf44f81638d}}, use this to unset). - cookbook ran by andrew@buster * 04:41 wm-bot2: Drained 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:40 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:40 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:39 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: cd2424d3-c9c9-4837-93b1-{{Gerrit|a6d8271e7873}}, use this to unset). - cookbook ran by andrew@buster * 04:39 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:39 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:38 wm-bot2: Safe reboot of 'cloudvirt1034.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 04:38 wm-bot2: Unset cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:38 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 1e85ec01-81c4-4664-b14f-{{Gerrit|b1cf6417988c}}, use this to unset). - cookbook ran by andrew@buster * 04:38 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: f498aaa4-fad3-42f8-92f4-{{Gerrit|48a98c108456}}, use this to unset). - cookbook ran by andrew@buster * 04:37 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:37 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:37 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:37 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:35 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 0ed281f3-d9b2-4c0a-bf88-{{Gerrit|db42fe848b2f}}, use this to unset). - cookbook ran by andrew@buster * 04:35 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: f417dffc-fa1b-4e65-938a-{{Gerrit|49e54b6fac6d}}, use this to unset). - cookbook ran by andrew@buster * 04:34 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:34 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:34 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:34 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:34 wm-bot2: Drained 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:33 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 923926e0-7194-4e13-98c4-{{Gerrit|a4f703b2907b}}, use this to unset). - cookbook ran by andrew@buster * 04:32 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:32 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:29 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: ca8f729b-92ad-45ec-b5da-{{Gerrit|55987b0fc9c2}}, use this to unset). - cookbook ran by andrew@buster * 04:29 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:29 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:25 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: abe6992d-dfba-4e50-993d-{{Gerrit|de0a551a177a}}, use this to unset). - cookbook ran by andrew@buster * 04:24 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:24 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:23 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: df91107d-bf34-49db-9756-{{Gerrit|44ad06162683}}, use this to unset). - cookbook ran by andrew@buster * 04:23 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: f45ec0b1-83c9-49e0-a3a8-{{Gerrit|897d5a48414f}}, use this to unset). - cookbook ran by andrew@buster * 04:23 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:22 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:22 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:22 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:18 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: 504ae503-bfad-41bd-9079-{{Gerrit|fb53a9c82a62}}, use this to unset). - cookbook ran by andrew@buster * 04:17 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:17 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:15 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: 8410faf9-667d-443d-bd17-{{Gerrit|bb3ca8e7f725}}, use this to unset). - cookbook ran by andrew@buster * 04:15 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:15 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:11 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 40441a37-0dd7-44a8-960b-{{Gerrit|f74f98f53618}}, use this to unset). - cookbook ran by andrew@buster * 04:10 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:10 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:08 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: 517cff2b-a643-4714-a31f-{{Gerrit|6bbe0767656a}}, use this to unset). - cookbook ran by andrew@buster * 04:08 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: 50b63ad6-4d8b-4fcf-9c17-{{Gerrit|27a106b6ec52}}, use this to unset). - cookbook ran by andrew@buster * 04:07 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:07 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:07 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:07 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:06 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: d912bdf5-54c7-490a-bdbe-{{Gerrit|ba9a23f07f4b}}, use this to unset). - cookbook ran by andrew@buster * 04:06 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: e75f16d0-e42f-4d72-a3bd-{{Gerrit|be0cbc5f200a}}, use this to unset). - cookbook ran by andrew@buster * 04:06 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: 87f6c417-3703-4f73-9876-{{Gerrit|20f6767dc8e1}}, use this to unset). - cookbook ran by andrew@buster * 04:05 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:05 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:05 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:05 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:05 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:05 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:02 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance (downtime id: be14f869-fdb0-4c38-84ff-{{Gerrit|3d1e00d227eb}}, use this to unset). - cookbook ran by andrew@buster * 04:02 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance (downtime id: feb93ea7-2c64-4831-b140-{{Gerrit|dad8311dbcef}}, use this to unset). - cookbook ran by andrew@buster * 04:02 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance (downtime id: c65a23e3-a809-4a75-88e4-{{Gerrit|29a8b351ecdb}}, use this to unset). - cookbook ran by andrew@buster * 04:02 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:01 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:01 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:01 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:01 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:01 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:00 wm-bot2: Safe reboot of 'cloudvirt1030.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 04:00 wm-bot2: Unset cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:57 wm-bot2: Drained 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:56 wm-bot2: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance (downtime id: 8c8f64ca-a4e8-47cf-a2ce-{{Gerrit|23059109fb6a}}, use this to unset). - cookbook ran by andrew@buster * 03:55 wm-bot2: Draining 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:55 wm-bot2: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:54 wm-bot2: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance (downtime id: 9d0457e8-a861-46ce-ac35-{{Gerrit|53975a46018e}}, use this to unset). - cookbook ran by andrew@buster * 03:53 wm-bot2: Draining 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:53 wm-bot2: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:44 wm-bot2: Safe reboot of 'cloudvirt1031.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 03:44 wm-bot2: Unset cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:42 wm-bot2: Safe reboot of 'cloudvirt1032.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 03:42 wm-bot2: Unset cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:40 wm-bot2: Drained 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:38 wm-bot2: Drained 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:30 wm-bot2: Set cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance (downtime id: 851d3b34-68f7-4c57-920b-{{Gerrit|2a64a9ea573d}}, use this to unset). - cookbook ran by andrew@buster * 03:29 wm-bot2: Draining 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:29 wm-bot2: Safe rebooting 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:17 wm-bot2: Set cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance (downtime id: 1060ab82-d53a-4d48-8e15-{{Gerrit|3198cabcfc2f}}, use this to unset). - cookbook ran by andrew@buster * 03:17 wm-bot2: Draining 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:17 wm-bot2: Safe rebooting 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:14 wm-bot2: Safe reboot of 'cloudvirt1029.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 03:14 wm-bot2: Unset cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:14 wm-bot2: Set cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance (downtime id: dde1dc09-44b5-42af-ac49-{{Gerrit|c58dbbae7cb4}}, use this to unset). - cookbook ran by andrew@buster * 03:14 wm-bot2: Draining 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:14 wm-bot2: Safe rebooting 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:13 wm-bot2: Drained 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:13 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: b59fa99b-8713-4e63-8c56-{{Gerrit|08bd71fdfbe3}}, use this to unset). - cookbook ran by andrew@buster * 03:13 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:13 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:07 wm-bot2: Set cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance (downtime id: 3547c843-6b15-4151-b452-{{Gerrit|6f618a886b7b}}, use this to unset). - cookbook ran by andrew@buster * 03:07 wm-bot2: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance (downtime id: 6f5808d0-afb2-4222-ae8e-{{Gerrit|4d953d49cd63}}, use this to unset). - cookbook ran by andrew@buster * 03:06 wm-bot2: Draining 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:06 wm-bot2: Safe rebooting 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:06 wm-bot2: Draining 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:06 wm-bot2: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:06 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: da738ba3-883f-4d76-bb33-{{Gerrit|7eb551de5fc2}}, use this to unset). - cookbook ran by andrew@buster * 03:05 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:05 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:04 wm-bot2: Safe reboot of 'cloudvirt1026.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 03:04 wm-bot2: Unset cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:03 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: f90ca4c1-78af-46bd-9262-{{Gerrit|0fa4edc2898e}}, use this to unset). - cookbook ran by andrew@buster * 03:02 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:02 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:02 wm-bot2: Safe reboot of 'cloudvirt1027.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 03:02 wm-bot2: Unset cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:00 wm-bot2: Drained 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:00 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance (downtime id: 62383663-9ae6-4a15-a2a6-{{Gerrit|62be06adbc96}}, use this to unset). - cookbook ran by andrew@buster * 02:59 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:59 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:59 wm-bot2: Drained 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:57 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance (downtime id: 6e0da344-49ad-451f-8b8d-{{Gerrit|ea8f853a6f89}}, use this to unset). - cookbook ran by andrew@buster * 02:57 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:57 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:52 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: c9ec5b7c-aa9f-4afd-a417-{{Gerrit|54c52c07ca47}}, use this to unset). - cookbook ran by andrew@buster * 02:51 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:51 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:48 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance (downtime id: 5509c648-58f6-49fd-90e8-{{Gerrit|ac9aa66e8b04}}, use this to unset). - cookbook ran by andrew@buster * 02:48 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:48 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:46 wm-bot2: Safe reboot of 'cloudvirt1024.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 02:46 wm-bot2: Unset cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:46 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance (downtime id: af4ba9ad-e765-4087-86aa-{{Gerrit|9111c0821175}}, use this to unset). - cookbook ran by andrew@buster * 02:46 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:45 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:45 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance (downtime id: 6da98123-8838-4d2e-8659-{{Gerrit|d769bfa292fe}}, use this to unset). - cookbook ran by andrew@buster * 02:44 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:44 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:43 wm-bot2: Safe reboot of 'cloudvirt1025.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 02:43 wm-bot2: Unset cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:42 wm-bot2: Drained 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:42 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 91860079-8fff-40db-9de7-{{Gerrit|17741d528ad6}}, use this to unset). - cookbook ran by andrew@buster * 02:41 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:41 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:39 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance (downtime id: a6a0f414-d305-4a09-87d2-{{Gerrit|066eec83642b}}, use this to unset). - cookbook ran by andrew@buster * 02:39 wm-bot2: Drained 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:38 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:38 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:38 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 6000898c-4e5e-49c8-8841-{{Gerrit|3262d6da0366}}, use this to unset). - cookbook ran by andrew@buster * 02:37 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:37 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:34 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 24a8b50e-5c04-40e6-a732-{{Gerrit|f0232b83e240}}, use this to unset). - cookbook ran by andrew@buster * 02:34 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 8d259aca-1803-4346-b5a4-{{Gerrit|a447673b3537}}, use this to unset). - cookbook ran by andrew@buster * 02:34 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:34 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:34 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:34 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:33 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 704aca7f-0069-4a4a-b8de-{{Gerrit|1aba7b9ca7ef}}, use this to unset). - cookbook ran by andrew@buster * 02:32 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 04c0fba3-e2cb-40b4-abfd-{{Gerrit|d24165c9bd55}}, use this to unset). - cookbook ran by andrew@buster * 02:32 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:32 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:32 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:32 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:28 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 25e5e0c8-5c27-460e-9b44-{{Gerrit|1cbd1c2fb551}}, use this to unset). - cookbook ran by andrew@buster * 02:28 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:28 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:25 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 5e62968a-0bb6-43b1-ae5e-{{Gerrit|e352442ef976}}, use this to unset). - cookbook ran by andrew@buster * 02:24 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:24 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:21 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 5b1a8e8f-8699-433d-b764-{{Gerrit|197862dc5cf1}}, use this to unset). - cookbook ran by andrew@buster * 02:20 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:20 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:15 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:15 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:13 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 299c466e-602b-4a8e-bb37-{{Gerrit|156553dbe426}}, use this to unset). - cookbook ran by andrew@buster * 02:13 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:13 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:12 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:12 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:11 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:11 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:03 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 6a64d619-c484-411f-ad9d-{{Gerrit|89d43b063737}}, use this to unset). - cookbook ran by andrew@buster * 02:02 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:02 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:46 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 301309a2-a8b5-4698-98d6-{{Gerrit|bb4aa0a75e45}}, use this to unset). - cookbook ran by andrew@buster * 00:46 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: b2f8c6a2-7224-4fb1-8e12-{{Gerrit|dff6dda02380}}, use this to unset). - cookbook ran by andrew@buster * 00:46 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:46 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:46 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:46 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:40 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:40 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:39 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:39 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:38 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: c0f62e0e-7865-4e77-a238-{{Gerrit|d707ceed53a8}}, use this to unset). - cookbook ran by andrew@buster * 00:38 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:38 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:36 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 96a368d5-232f-4187-b0bb-{{Gerrit|5923a50fff71}}, use this to unset). - cookbook ran by andrew@buster * 00:36 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:36 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:35 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 31b6bfcb-81cb-4d89-b977-{{Gerrit|190899ea2e64}}, use this to unset). - cookbook ran by andrew@buster * 00:35 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:35 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:28 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 50f763d4-36ea-4241-8e71-{{Gerrit|cc6aa20ccc9e}}, use this to unset). - cookbook ran by andrew@buster * 00:28 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: cd3d17af-ee15-4a9f-8750-{{Gerrit|66b97a46ea12}}, use this to unset). - cookbook ran by andrew@buster * 00:27 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:27 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:27 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:27 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:22 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:22 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:22 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 9b075166-d424-42c0-8e06-{{Gerrit|f74a47abb798}}, use this to unset). - cookbook ran by andrew@buster * 00:21 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:21 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:21 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:21 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster === 2022-07-18 === * 23:11 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: e0aab64f-d911-4d7e-9a97-{{Gerrit|c59d9868ceea}}, use this to unset). - cookbook ran by andrew@buster * 23:11 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 447d9c7a-a00e-41db-93ed-{{Gerrit|6d2bfc66e5df}}, use this to unset). - cookbook ran by andrew@buster * 23:11 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:11 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:11 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:11 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:01 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 46c2ddf2-50c3-4824-ba72-{{Gerrit|d7fb3fa4c5e0}}, use this to unset). - cookbook ran by andrew@buster * 23:01 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 6facdf4f-7bb8-42cd-8ac8-{{Gerrit|b4412fc4cf70}}, use this to unset). - cookbook ran by andrew@buster * 23:00 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:00 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:00 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:00 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:52 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance (downtime id: 23190573-fc0d-4872-8343-{{Gerrit|e35b84132028}}, use this to unset). - cookbook ran by andrew@buster * 22:51 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:51 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:51 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 990313c4-5fa6-4de7-a8bf-{{Gerrit|b361f0c79e90}}, use this to unset). - cookbook ran by andrew@buster * 22:50 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:50 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:48 wm-bot2: Safe reboot of 'cloudvirt1022.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 22:48 wm-bot2: Unset cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 22:42 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: 59367d39-bd2f-47fd-becd-{{Gerrit|255bf823ff12}}, use this to unset). - cookbook ran by andrew@buster * 22:42 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:42 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:40 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 38602815-d187-4455-9676-{{Gerrit|59607505bba7}}, use this to unset). - cookbook ran by andrew@buster * 22:39 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: 00f21e1b-d014-4bc5-995f-{{Gerrit|e25201fc77c4}}, use this to unset). - cookbook ran by andrew@buster * 22:39 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:39 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:39 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:39 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:31 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: d05d07b2-8c60-46de-a89e-{{Gerrit|76987901901b}}, use this to unset). - cookbook ran by andrew@buster * 22:31 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance (downtime id: 47f6b9e6-ceab-4fb6-8f76-{{Gerrit|0de4341d0841}}, use this to unset). - cookbook ran by andrew@buster * 22:31 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:31 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:30 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:30 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:30 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: 599042a5-da25-4985-b502-{{Gerrit|328fe40a72d5}}, use this to unset). - cookbook ran by andrew@buster * 22:29 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:29 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:29 wm-bot2: Drained 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:28 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: a554d107-c14d-41f8-b180-{{Gerrit|093e2cf73e72}}, use this to unset). - cookbook ran by andrew@buster * 22:28 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: 9fef70c7-6f7d-4b33-a9bc-{{Gerrit|8c39630dc182}}, use this to unset). - cookbook ran by andrew@buster * 22:26 wm-bot2: Drained 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:02 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: 3744c0fa-084d-4198-98ca-{{Gerrit|66c7a5210f41}}, use this to unset). - cookbook ran by andrew@buster * 22:02 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: 01caf1e1-df64-474a-94d5-{{Gerrit|f37751e23d62}}, use this to unset). - cookbook ran by andrew@buster * 22:01 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:01 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:01 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:01 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:59 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: 3018c2fc-9cd5-45c8-a160-{{Gerrit|d5ad5728c8d5}}, use this to unset). - cookbook ran by andrew@buster * 21:59 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:59 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:02 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: d760f916-32ea-4194-8974-{{Gerrit|f36064a9866c}}, use this to unset). - cookbook ran by andrew@buster * 21:02 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:02 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:00 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: cb2a38f3-e9ca-4bf6-989c-{{Gerrit|a27be2b2d88c}}, use this to unset). - cookbook ran by andrew@buster * 20:59 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:59 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:58 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: 93fa9477-54b8-46e7-8508-{{Gerrit|a2d492528d1f}}, use this to unset). - cookbook ran by andrew@buster * 20:57 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:57 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:52 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: f6cf5b29-ce8b-426e-af1c-{{Gerrit|cab74e0c4a3a}}, use this to unset). - cookbook ran by andrew@buster * 20:52 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: fe26e612-d76e-4120-b60a-{{Gerrit|4f300add224f}}, use this to unset). - cookbook ran by andrew@buster * 20:52 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: 2abd94be-b6d8-4042-a9ce-{{Gerrit|998e41d932f4}}, use this to unset). - cookbook ran by andrew@buster * 20:51 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:51 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:51 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:51 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:51 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:51 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:34 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:34 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:28 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: 10973c71-30d8-44bf-a310-{{Gerrit|d0c1af76d399}}, use this to unset). - cookbook ran by andrew@buster * 20:28 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: f8f1b214-ec76-4157-a8e6-{{Gerrit|6374e5d14608}}, use this to unset). - cookbook ran by andrew@buster * 20:27 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:27 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:27 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:27 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:11 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: 95dae613-2b04-472e-a8f9-{{Gerrit|e71c3fbb8d0e}}, use this to unset). - cookbook ran by andrew@buster * 20:10 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:10 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:03 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance (downtime id: c804f4db-ea63-44e9-aff8-{{Gerrit|fea54c238345}}, use this to unset). - cookbook ran by andrew@buster * 20:03 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance (downtime id: 597a901d-4624-43de-b986-{{Gerrit|c4c6c2fcf80c}}, use this to unset). - cookbook ran by andrew@buster * 20:03 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:03 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:02 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance (downtime id: fdceed0d-6fa0-4806-8fb6-{{Gerrit|3b321a5a1623}}, use this to unset). - cookbook ran by andrew@buster * 20:02 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:02 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:02 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:02 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:01 wm-bot2: Safe reboot of 'cloudvirt1017.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 20:01 wm-bot2: Unset cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:57 wm-bot2: Drained 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:37 wm-bot2: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance (downtime id: b0904345-9aa3-4202-b992-{{Gerrit|78644141517c}}, use this to unset). - cookbook ran by andrew@buster * 19:36 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:36 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:31 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:31 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:30 wm-bot2: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:30 wm-bot2: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster === 2022-07-13 === * 14:48 bd808: Added Tavvi as member of #acl*wmcs-team === 2022-07-12 === * 12:04 wm-bot2: Ceph cluster at <nowiki>{</nowiki>self.deployment<nowiki>}</nowiki> set out of maintenance. - cookbook ran by dcaro@vulcanus * 12:03 wm-bot2: Set the ceph cluster for codfw1dev in maintenance, alert silence ids: db32805f-a033-4e0e-8fc7-{{Gerrit|ed0c2e9f8be1}},f4e698f0-4b51-4b07-9ba8-{{Gerrit|0b296ca1c4fb}},39ad5325-44ed-44d1-bba5-{{Gerrit|91a85c7a8401}},a584a3c6-9dae-41e9-ac82-{{Gerrit|a91cd35af8fb}} - cookbook ran by dcaro@vulcanus * 08:37 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2005-dev', 'cloudnet2006-dev'] - cookbook ran by dcaro@vulcanus * 08:37 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:31 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:31 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:25 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:25 wm-bot2: Rebooting all the cloudnet nodes cloudnet2005-dev,cloudnet2006-dev - cookbook ran by dcaro@vulcanus === 2022-07-08 === * 15:57 wm-bot2: Finished rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 15:52 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 15:52 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus === 2022-07-07 === * 07:17 wm-bot2: Finished rebooting node cloudcephosd1015.eqiad.wmnet ([[phab:T312509|T312509]]) - cookbook ran by dcaro@vulcanus * 07:12 wm-bot2: Rebooting node cloudcephosd1015.eqiad.wmnet ([[phab:T312509|T312509]]) - cookbook ran by dcaro@vulcanus === 2022-07-06 === * 17:50 wm-bot2: Set the ceph cluster for eqiad1 in maintenance, alert silence ids: ['8a5b9eee-48c0-474d-8277-faeb05a2ea61', '65aad0fc-d887-47a3-b20c-d1ed461a2411', '86b078ae-3a27-4063-8c7c-198a2fe0c172'] - cookbook ran by dcaro@vulcanus === 2022-07-04 === * 13:27 wm-bot2: Rebooting cloudgw host cloudgw1002.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:27 wm-bot2: Rebooted cloudgw host cloudgw1001.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:23 wm-bot2: Rebooting cloudgw host cloudgw1001.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:23 wm-bot2: Rebooting all the cloudgw nodes from the eqiad1 deployment: cloudgw1001.eqiad.wmnet,cloudgw1002.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:13 wm-bot2: Finished rebooting the cloudgw nodes ['cloudgw2001-dev.codfw.wmnet', 'cloudgw2002-dev.codfw.wmnet', 'cloudgw2003-dev.codfw.wmnet'] - cookbook ran by dcaro@vulcanus * 13:12 wm-bot2: Rebooted cloudgw host cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:09 wm-bot2: Rebooting cloudgw host cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:08 wm-bot2: Rebooted cloudgw host cloudgw2002-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:05 wm-bot2: Rebooting cloudgw host cloudgw2002-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:05 wm-bot2: Rebooted cloudgw host cloudgw2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:01 wm-bot2: Rebooting cloudgw host cloudgw2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:01 wm-bot2: Rebooting all the cloudgw nodes cloudgw2001-dev.codfw.wmnet,cloudgw2002-dev.codfw.wmnet,cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 12:56 wm-bot2: Rebooting all the cloudgw nodes cloudgw2001-dev.codfw.wmnet,cloudgw2002-dev.codfw.wmnet,cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 12:52 wm-bot2: Rebooting all the cloudgw nodes cloudgw2001-dev.codfw.wmnet,cloudgw2002-dev.codfw.wmnet,cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:56 wm-bot2: Finished rebooting the cloudgw nodes ['cloudgw2001-dev.codfw.wmnet', 'cloudgw2002-dev.codfw.wmnet', 'cloudgw2003-dev.codfw.wmnet'] - cookbook ran by dcaro@vulcanus * 10:56 wm-bot2: Rebooted cloudgw host cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:51 wm-bot2: Rebooting cloudgw host cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:51 wm-bot2: Rebooted cloudgw host cloudgw2002-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:47 wm-bot2: Rebooting cloudgw host cloudgw2002-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:47 wm-bot2: Rebooted cloudgw host cloudgw2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:42 wm-bot2: Rebooting cloudgw host cloudgw2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:42 wm-bot2: Rebooting all the cloudgw nodes cloudgw2001-dev.codfw.wmnet,cloudgw2002-dev.codfw.wmnet,cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:41 wm-bot2: Rebooting all the cloudgw nodes cloudgw2001-dev.codfw.wmnet,cloudgw2002-dev.codfw.wmnet,cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 10:40 wm-bot2: Rebooting all the cloudgw nodes cloudgw2001-dev.codfw.wmnet,cloudgw2002-dev.codfw.wmnet,cloudgw2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 09:07 wm-bot2: Rebooting cloudnet host cloudnet1004.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 09:06 wm-bot2: Rebooted cloudnet host cloudnet1003.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 09:00 wm-bot2: Rebooting cloudnet host cloudnet1003.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 09:00 wm-bot2: Rebooting all the cloudnet nodes cloudnet1003,cloudnet1004 - cookbook ran by dcaro@vulcanus * 08:59 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2006-dev', 'cloudnet2005-dev'] - cookbook ran by dcaro@vulcanus * 08:59 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:53 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:53 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:47 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:47 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 08:32 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2006-dev', 'cloudnet2005-dev'] - cookbook ran by dcaro@vulcanus * 08:32 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:25 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:25 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:19 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 08:19 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 07:55 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 07:51 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 07:44 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2006-dev', 'cloudnet2005-dev'] - cookbook ran by dcaro@vulcanus * 07:44 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:40 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:40 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:36 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:36 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 07:09 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:05 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:05 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 07:04 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus === 2022-07-03 === * 21:27 andrewbogott: rebuilding rabbit cluster in codfw1dev to get rid of some queues so unresponsive that they can't otherwise be deleted === 2022-07-02 === * 11:05 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus === 2022-07-01 === * 15:35 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2006-dev', 'cloudnet2005-dev'] - cookbook ran by dcaro@vulcanus * 15:35 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:30 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:29 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:25 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:25 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 15:13 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2006-dev', 'cloudnet2005-dev'] - cookbook ran by dcaro@vulcanus * 15:13 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:08 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:07 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:03 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 15:03 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 14:42 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2006-dev', 'cloudnet2005-dev'] - cookbook ran by dcaro@vulcanus * 14:42 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:37 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:36 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:32 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:31 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 14:23 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:19 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:19 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 14:19 wm-bot2: Finished rebooting the cloudnet nodes ['cloudnet2006-dev', 'cloudnet2005-dev'] - cookbook ran by dcaro@vulcanus * 14:09 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:06 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:02 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:58 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:58 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 13:55 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:54 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 13:52 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:51 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 13:37 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus * 13:34 wm-bot2: Rebooted cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:31 wm-bot2: Rebooting cloudnet host cloudnet2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 13:30 wm-bot2: Rebooting all the cloudnet nodes cloudnet2006-dev,cloudnet2005-dev - cookbook ran by dcaro@vulcanus === 2022-06-30 === * 18:17 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 18:13 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:45 wm-bot2: Rebooted cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 07:40 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 06:52 wm-bot2: Rebooting cloudnet host cloudnet2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus === 2022-06-29 === * 08:45 wm-bot2: Finished rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 08:41 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus === 2022-06-28 === * 13:03 taavi: grant the tools project access to the g3.cores16.ram64.disk20.10xiops flavor [[phab:T301949|T301949]] === 2022-06-22 === * 16:48 taavi: restart designate-*.service on both cloudservices nodes * 16:44 taavi: restart nova-conductor on all the cloudcontrol nodes * 12:50 andrewbogott: rebooting each eqiad1 cloudcontrol node in hopes of getting a baseline re: openstack instability === 2022-06-21 === * 04:48 andrewbogott: stopping nova-fullstack agent on cloudcontrol1003; it's going to page us otherwise and we're all AFK tomorrow * 04:02 andrewbogott: restarting rabbitmq on cloudcontrol100x (one at a time) === 2022-06-17 === * 17:53 andrewbogott: switching to a new python-based health check for galera and haproxy. This may make things more stable, or it may not. [[phab:T310664|T310664]] * 06:15 taavi: restart neutron-linuxbridge-agent on cloudvirt1046 === 2022-06-15 === * 11:34 taavi: restart neutron-linuxbridge-agent on cloudvirt1022 === 2022-06-14 === * 16:26 wm-bot2: OSDs (['cloudcephosd1001', 'cloudcephosd1002', 'cloudcephosd1003', 'cloudcephosd1004', 'cloudcephosd1005', 'cloudcephosd1006', 'cloudcephosd1007', 'cloudcephosd1008', 'cloudcephosd1009', 'cloudcephosd1010', 'cloudcephosd1011', 'cloudcephosd1012', 'cloudcephosd1013', 'cloudcephosd1014', 'cloudcephosd1015', 'cloudcephosd1016', 'cloudcephosd1017', 'cloudcephosd1018', 'cloudcephosd1019', 'cloudcephosd1020', 'cloudcephosd1021', 'cl * 14:38 wm-bot2: Upgrading OSDs and rebooting the nodes ['cloudcephosd1001', 'cloudcephosd1002', 'cloudcephosd1003', 'cloudcephosd1004', 'cloudcephosd1005', 'cloudcephosd1006', 'cloudcephosd1007', 'cloudcephosd1008', 'cloudcephosd1009', 'cloudcephosd1010', 'cloudcephosd1011', 'cloudcephosd1012', 'cloudcephosd1013', 'cloudcephosd1014', 'cloudcephosd1015', 'cloudcephosd1016', 'cloudcephosd1017', 'cloudcephosd1018', 'cloudcephosd1019', 'cloudceph * 12:55 wm-bot2: OSDs (['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev']) upgraded successfully B-) ([[phab:T309786|T309786]]) - cookbook ran by dcaro@vulcanus * 12:44 wm-bot2: Upgrading OSDs and rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T309786|T309786]]) - cookbook ran by dcaro@vulcanus === 2022-06-13 === * 11:14 wm-bot2: Finished rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T309789|T309789]]) - cookbook ran by dcaro@vulcanus * 11:08 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T309789|T309789]]) - cookbook ran by dcaro@vulcanus * 11:07 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T309789|T309789]]) - cookbook ran by dcaro@vulcanus * 11:05 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T309789|T309789]]) - cookbook ran by dcaro@vulcanus * 11:04 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T309789|T309789]]) - cookbook ran by dcaro@vulcanus * 11:03 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T309789|T309789]]) - cookbook ran by dcaro@vulcanus * 09:15 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet ([[phab:T309789|T309789]]) - cookbook ran by dcaro@vulcanus === 2022-06-06 === * 13:21 andrewbogott: restarting mysql/galera on cloudcontrol100x in an attempt to stabilize some flapping there === 2022-06-02 === * 07:51 taavi: restart neutron-linuxbridge-agent.service on cloudvirt1034 [[phab:T309732|T309732]] * 00:37 andrewbogott: updated nameservers for codfw1dev instances via 'openstack subnet set --dns-nameserver etc.' === 2022-06-01 === * 17:11 andrewbogott: restarting designate services in cloudservices1xxx hosts * 16:37 taavi: root@cloudcontrol1005:~# cinder reset-state --state available 491066a4-16a9-4ce4-a9f6-{{Gerrit|182660616d77}} for [[phab:T309659|T309659]] === 2022-05-30 === * 17:43 andrewbogott: restarting neutron-rpc and neutron-api services on cloudcontrol1xxx === 2022-05-29 === * 14:55 andrewbogott: restarting nova services on all eqiad1 cloudcontrol nodes to recover from rabbit breakage * 14:15 andrewbogott: restarting rabbitmq on all eqiad1 cloudcontrol nodes (one at a time) === 2022-05-25 === * 20:01 balloons: clean up cinder backup a bit, restart service due to network outage * 20:01 balloons: cloudvirt1029 restarted nova due to network outage === 2022-05-19 === * 15:21 andrewbogott: resetting password for the 'troveguest' rabbitmq user. I think I may have broken this during a recent rebuild of the rabbitmq cluster === 2022-05-18 === * 15:42 andrewbogott: updated the 'debian-11.0-bullseye' glance image with a fresh build === 2022-05-14 === * 11:33 taavi: deleted projects 'ores' and 'ores-staging' [[phab:T308102|T308102]] === 2022-05-13 === * 06:20 wm-bot2: Safe reboot of 'cloudvirt1045.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 06:20 wm-bot2: Unset cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 06:16 wm-bot2: Drained 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 06:16 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 06:15 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 06:15 wm-bot2: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 06:11 wm-bot2: Safe reboot of 'cloudvirt1044.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 06:11 wm-bot2: Unset cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 06:10 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 06:10 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 06:09 wm-bot2: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 06:07 wm-bot2: Drained 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 06:06 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 06:06 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 06:05 wm-bot2: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:51 wm-bot2: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:50 wm-bot2: Draining 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:50 wm-bot2: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:49 wm-bot2: Set cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:49 wm-bot2: Safe reboot of 'cloudvirt1043.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 05:49 wm-bot2: Unset cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:49 wm-bot2: Draining 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:49 wm-bot2: Safe rebooting 'cloudvirt1044.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:47 wm-bot2: Safe reboot of 'cloudvirt1042.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 05:47 wm-bot2: Unset cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:45 wm-bot2: Drained 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:45 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:45 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:44 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:44 wm-bot2: Drained 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:42 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:42 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:42 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:41 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:40 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:40 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:38 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:37 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:37 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:30 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:29 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:29 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:19 wm-bot2: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:18 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:18 wm-bot2: Draining 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:18 wm-bot2: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:18 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:18 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:12 wm-bot2: Safe reboot of 'cloudvirt1040.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 05:12 wm-bot2: Unset cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:08 wm-bot2: Drained 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:02 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:02 wm-bot2: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 05:02 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:02 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:02 wm-bot2: Draining 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 05:01 wm-bot2: Safe rebooting 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:52 wm-bot2: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:51 wm-bot2: Draining 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:51 wm-bot2: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:48 wm-bot2: Safe reboot of 'cloudvirt1041.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 04:48 wm-bot2: Unset cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:44 wm-bot2: Drained 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:31 wm-bot2: Set cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:30 wm-bot2: Draining 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:30 wm-bot2: Safe rebooting 'cloudvirt1041.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:30 wm-bot2: Safe reboot of 'cloudvirt1039.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 04:30 wm-bot2: Unset cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:27 wm-bot2: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:26 wm-bot2: Draining 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:26 wm-bot2: Safe rebooting 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:26 wm-bot2: Drained 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:26 wm-bot2: Set cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:25 wm-bot2: Draining 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:25 wm-bot2: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:24 wm-bot2: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:23 wm-bot2: Draining 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:23 wm-bot2: Safe rebooting 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:23 wm-bot2: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:22 wm-bot2: Draining 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:22 wm-bot2: Safe rebooting 'cloudvirt1040.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:21 wm-bot2: Safe reboot of 'cloudvirt1038.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 04:21 wm-bot2: Unset cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:18 wm-bot2: Drained 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:16 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:16 wm-bot2: Set cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:16 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:16 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:15 wm-bot2: Draining 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:15 wm-bot2: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:37 wm-bot2: Set cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:36 wm-bot2: Draining 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:36 wm-bot2: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:34 wm-bot2: Safe reboot of 'cloudvirt1037.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 03:34 wm-bot2: Unset cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:27 wm-bot2: Drained 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:27 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:26 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:26 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:26 wm-bot2: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:25 wm-bot2: Draining 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:25 wm-bot2: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:22 wm-bot2: Unset cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:55 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:55 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:54 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:54 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:54 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:54 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:05 wm-bot2: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:05 wm-bot2: Draining 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:05 wm-bot2: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:04 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:04 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:03 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 01:23 wm-bot2: Safe reboot of 'cloudvirt1035.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 01:23 wm-bot2: Unset cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 01:19 wm-bot2: Drained 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 01:01 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 01:01 wm-bot2: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 01:00 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 01:00 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 01:00 wm-bot2: Draining 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 01:00 wm-bot2: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:25 wm-bot2: Safe reboot of 'cloudvirt1033.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 00:25 wm-bot2: Unset cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:21 wm-bot2: Drained 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:20 wm-bot2: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:19 wm-bot2: Draining 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:19 wm-bot2: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. - cookbook ran by andrew@buster * 00:11 wm-bot2: Safe reboot of 'cloudvirt1034.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 00:11 wm-bot2: Unset cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:07 wm-bot2: Drained 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster === 2022-05-12 === * 23:55 wm-bot2: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 23:55 wm-bot2: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 23:54 wm-bot2: Draining 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:54 wm-bot2: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:54 wm-bot2: Draining 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 23:54 wm-bot2: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:23 wm-bot2: Safe reboot of 'cloudvirt1031.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 22:23 wm-bot2: Unset cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 22:20 wm-bot2: Drained 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 22:17 wm-bot2: Safe reboot of 'cloudvirt1032.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 22:17 wm-bot2: Unset cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 22:13 wm-bot2: Drained 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:57 wm-bot2: Set cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:56 wm-bot2: Draining 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:56 wm-bot2: Safe rebooting 'cloudvirt1032.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:55 wm-bot2: Draining 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:55 wm-bot2: Safe rebooting 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:54 wm-bot2: Safe reboot of 'cloudvirt1030.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:54 wm-bot2: Unset cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:53 wm-bot2: Set cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:52 wm-bot2: Draining 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:52 wm-bot2: Safe rebooting 'cloudvirt1031.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:51 wm-bot2: Drained 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:44 wm-bot2: Safe reboot of 'cloudvirt1029.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:44 wm-bot2: Unset cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:42 wm-bot2: Drained 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:36 wm-bot2: Safe reboot of 'cloudvirt-wdqs1001.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:36 wm-bot2: Unset cloudvirt 'cloudvirt-wdqs1001.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:33 wm-bot2: Drained 'cloudvirt-wdqs1001.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:33 wm-bot2: Set cloudvirt 'cloudvirt-wdqs1001.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:32 wm-bot2: Draining 'cloudvirt-wdqs1001.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:32 wm-bot2: Safe rebooting 'cloudvirt-wdqs1001.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:32 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:31 wm-bot2: Safe reboot of 'cloudvirt-wdqs1002.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:31 wm-bot2: Unset cloudvirt 'cloudvirt-wdqs1002.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:31 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:31 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:30 wm-bot2: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:29 wm-bot2: Draining 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:29 wm-bot2: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:29 wm-bot2: Drained 'cloudvirt-wdqs1002.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:28 wm-bot2: Set cloudvirt 'cloudvirt-wdqs1002.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:28 wm-bot2: Draining 'cloudvirt-wdqs1002.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:28 wm-bot2: Safe rebooting 'cloudvirt-wdqs1002.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:22 wm-bot2: Safe reboot of 'cloudvirt1026.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:22 wm-bot2: Unset cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:21 wm-bot2: Safe reboot of 'cloudvirt-wdqs1003.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:21 wm-bot2: Unset cloudvirt 'cloudvirt-wdqs1003.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:18 wm-bot2: Drained 'cloudvirt-wdqs1003.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:18 wm-bot2: Drained 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:18 wm-bot2: Set cloudvirt 'cloudvirt-wdqs1003.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:17 wm-bot2: Draining 'cloudvirt-wdqs1003.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:17 wm-bot2: Safe rebooting 'cloudvirt-wdqs1003.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:17 wm-bot2: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:16 wm-bot2: Draining 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:16 wm-bot2: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:14 wm-bot2: Safe reboot of 'cloudvirt1025.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:14 wm-bot2: Unset cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:11 wm-bot2: Safe reboot of 'cloudvirt1046.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 21:11 wm-bot2: Unset cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:10 wm-bot2: Drained 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:08 wm-bot2: Drained 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:08 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:07 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:07 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:05 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:04 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:04 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:00 wm-bot2: Set cloudvirt 'cloudvirt1046.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:59 wm-bot2: Draining 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:59 wm-bot2: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:59 wm-bot2: Safe reboot of 'cloudvirt1047.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 20:59 wm-bot2: Unset cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:57 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:57 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:57 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:55 wm-bot2: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:55 wm-bot2: Drained 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:54 wm-bot2: Draining 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:54 wm-bot2: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:54 wm-bot2: Safe reboot of 'cloudvirt1024.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 20:54 wm-bot2: Unset cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:53 wm-bot2: Set cloudvirt 'cloudvirt1047.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:52 wm-bot2: Draining 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:52 wm-bot2: Safe rebooting 'cloudvirt1047.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:50 wm-bot2: Drained 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:49 wm-bot2: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:49 wm-bot2: Draining 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:49 wm-bot2: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:48 wm-bot2: Safe reboot of 'cloudvirt1023.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 20:48 wm-bot2: Unset cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:44 wm-bot2: Drained 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:44 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:43 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:43 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:34 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:34 wm-bot2: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:34 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:34 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:34 wm-bot2: Draining 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:34 wm-bot2: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:31 wm-bot2: Safe reboot of 'cloudvirt1027.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 20:31 wm-bot2: Unset cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:28 wm-bot2: Drained 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:28 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:27 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:27 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:23 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:22 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:22 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:11 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:10 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:10 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:07 wm-bot2: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:07 wm-bot2: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:06 wm-bot2: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:06 wm-bot2: Safe reboot of 'cloudvirt1022.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 20:05 wm-bot2: Unset cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:02 wm-bot2: Drained 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:01 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:00 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:00 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:58 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:57 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:57 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:36 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:35 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:35 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:06 andrewbogott: stopping nfs-server on labstore1004 in preparation for reboot * 04:12 andrewbogott: rebooting primary bastion (bastion-eqiad1-03.bastion.eqiad1.wikimedia.cloud) in hopes of resolving a problem with ssh proxying === 2022-05-11 === * 18:48 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 18:48 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:48 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:39 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 18:38 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:38 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:04 wm-bot2: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 18:03 wm-bot2: Draining 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 18:03 wm-bot2: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. - cookbook ran by andrew@buster * 08:56 wm-bot2: Finished rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 08:52 wm-bot2: Rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 07:53 dcaro: test * 04:28 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 04:27 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 04:27 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:44 wm-bot2: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:43 wm-bot2: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:43 wm-bot2: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:42 wm-bot2: Safe reboot of 'cloudvirt1021.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 03:42 wm-bot2: Unset cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:39 wm-bot2: Drained 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:23 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:22 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:22 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:09 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:08 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:08 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:04 andrewbogott: reset and recreated the rabbitmq cluster in eqiad1 to get around some broken queues. * 03:02 wm-bot2: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 03:01 wm-bot2: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 03:01 wm-bot2: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:28 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 02:25 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 02:25 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster === 2022-05-10 === * 21:43 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:40 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:40 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:35 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 21:32 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 21:32 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:05 wm-bot: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 20:02 wm-bot: Draining 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:01 wm-bot: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:00 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 20:00 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:57 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:57 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:55 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:55 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:47 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:46 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:46 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:45 wm-bot: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:45 wm-bot: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:44 wm-bot: Draining 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:44 wm-bot: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:40 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:39 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:39 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:37 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:36 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:36 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:33 wm-bot: Safe reboot of 'cloudvirt1017.eqiad.wmnet' finished successfully. - cookbook ran by andrew@buster * 19:33 wm-bot: Unset cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:29 wm-bot: Drained 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:06 wm-bot: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 19:05 wm-bot: Draining 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 19:05 wm-bot: Safe rebooting 'cloudvirt1017.eqiad.wmnet'. - cookbook ran by andrew@buster * 15:41 andrewbogott: rebooting cloud*-dev for [[phab:T307668|T307668]] * 13:59 taavi: manually attached [[User:Dreamy Jazz]] to wikitech for a password reset (https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin#Manually_associate_an_LDAP_account_with_wikitech) === 2022-05-07 === * 01:33 wm-bot: Drained 'cloudvirt1016.eqiad.wmnet'. - cookbook ran by andrew@buster * 01:32 wm-bot: Set cloudvirt 'cloudvirt1016.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 01:30 wm-bot: Draining 'cloudvirt1016.eqiad.wmnet'. - cookbook ran by andrew@buster * 01:21 wm-bot: Set cloudvirt 'cloudvirt1016.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 01:18 wm-bot: Draining 'cloudvirt1016.eqiad.wmnet'. - cookbook ran by andrew@buster === 2022-05-03 === * 20:38 andrewbogott: upgrading clouddb2001-dev in place * 18:18 taavi: updated 'puppet-enc' endpoints on the keystone catalog to use https and port 443 === 2022-05-02 === * 16:56 dcaro: rebooting cloudmetrics1001 === 2022-04-29 === * 14:22 andrewbogott: changing login.toolforge.org, bastion.toolforge.org, and dev.toolforge.org dns entries to refer to the new Buster bastions [[phab:T277653|T277653]] https://wikitech.wikimedia.org/wiki/News/Toolforge_Stretch_deprecation#Timeline === 2022-04-27 === * 14:51 wm-bot: Finished rebooting the nodes ['cloudcephosd1001', 'cloudcephosd1002', 'cloudcephosd1003', 'cloudcephosd1004', 'cloudcephosd1005', 'cloudcephosd1006', 'cloudcephosd1007', 'cloudcephosd1008', 'cloudcephosd1009', 'cloudcephosd1010', 'cloudcephosd1011', 'cloudcephosd1012', 'cloudcephosd1013', 'cloudcephosd1014', 'cloudcephosd1015', 'cloudcephosd1016', 'cloudcephosd1017', 'cloudcephosd1018', 'cloudcephosd1019', 'cloudcephosd1020', 'cloud * 14:50 wm-bot: Finished rebooting node cloudcephosd1024.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:46 wm-bot: Rebooting node cloudcephosd1024.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:46 wm-bot: Finished rebooting node cloudcephosd1023.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:41 wm-bot: Rebooting node cloudcephosd1023.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:41 wm-bot: Finished rebooting node cloudcephosd1022.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:35 wm-bot: Rebooting node cloudcephosd1022.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:35 wm-bot: Finished rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:31 wm-bot: Rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:31 wm-bot: Finished rebooting node cloudcephosd1020.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:27 wm-bot: Rebooting node cloudcephosd1020.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:27 wm-bot: Finished rebooting node cloudcephosd1019.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:23 wm-bot: Rebooting node cloudcephosd1019.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:23 wm-bot: Finished rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:13 wm-bot: Rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:13 wm-bot: Finished rebooting node cloudcephosd1017.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:09 wm-bot: Rebooting node cloudcephosd1017.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:09 wm-bot: Finished rebooting node cloudcephosd1016.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:05 wm-bot: Rebooting node cloudcephosd1016.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:05 wm-bot: Finished rebooting node cloudcephosd1015.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:01 wm-bot: Rebooting node cloudcephosd1015.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:01 wm-bot: Finished rebooting node cloudcephosd1014.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:57 wm-bot: Rebooting node cloudcephosd1014.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:57 wm-bot: Finished rebooting node cloudcephosd1013.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:44 wm-bot: Rebooting node cloudcephosd1013.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:43 wm-bot: Finished rebooting node cloudcephosd1012.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:39 wm-bot: Rebooting node cloudcephosd1012.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:39 wm-bot: Finished rebooting node cloudcephosd1011.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:35 wm-bot: Rebooting node cloudcephosd1011.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:35 wm-bot: Finished rebooting node cloudcephosd1010.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:31 wm-bot: Rebooting node cloudcephosd1010.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:31 wm-bot: Finished rebooting node cloudcephosd1009.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:26 wm-bot: Rebooting node cloudcephosd1009.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:26 wm-bot: Finished rebooting node cloudcephosd1008.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:14 wm-bot: Rebooting node cloudcephosd1008.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:14 wm-bot: Finished rebooting node cloudcephosd1007.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:10 wm-bot: Rebooting node cloudcephosd1007.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:10 wm-bot: Finished rebooting node cloudcephosd1006.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:05 wm-bot: Rebooting node cloudcephosd1006.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:05 wm-bot: Finished rebooting node cloudcephosd1005.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:01 wm-bot: Rebooting node cloudcephosd1005.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:01 wm-bot: Finished rebooting node cloudcephosd1004.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:57 wm-bot: Rebooting node cloudcephosd1004.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:57 wm-bot: Finished rebooting node cloudcephosd1003.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:53 wm-bot: Rebooting node cloudcephosd1003.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:53 wm-bot: Finished rebooting node cloudcephosd1002.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:50 wm-bot: Rebooting node cloudcephosd1002.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:50 wm-bot: Finished rebooting node cloudcephosd1001.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:46 wm-bot: Rebooting node cloudcephosd1001.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:46 wm-bot: Rebooting the nodes cloudcephosd1001,cloudcephosd1002,cloudcephosd1003,cloudcephosd1004,cloudcephosd1005,cloudcephosd1006,cloudcephosd1007,cloudcephosd1008,cloudcephosd1009,cloudcephosd1010,cloudcephosd1011,cloudcephosd1012,cloudcephosd1013,cloudcephosd1014,cloudcephosd1015,cloudcephosd1016,cloudcephosd1017,cloudcephosd1018,cloudcephosd1019,cloudcephosd1020,cloudcephosd1021,cloudcephosd1022,cloudcephosd1023,cloudcephosd1024 - cookbo * 12:15 wm-bot: Finished rebooting the nodes ['cloudcephmon1001', 'cloudcephmon1002', 'cloudcephmon1003'] - cookbook ran by dcaro@vulcanus * 12:15 wm-bot: Finished rebooting node cloudcephmon1003.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:12 wm-bot: Rebooting node cloudcephmon1003.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:12 wm-bot: Finished rebooting node cloudcephmon1002.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:09 wm-bot: Rebooting node cloudcephmon1002.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:09 wm-bot: Finished rebooting node cloudcephmon1001.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:07 wm-bot: Rebooting node cloudcephmon1001.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 12:07 wm-bot: Rebooting the nodes cloudcephmon1001,cloudcephmon1002,cloudcephmon1003 - cookbook ran by dcaro@vulcanus * 12:05 wm-bot: Finished rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] - cookbook ran by dcaro@vulcanus * 12:05 wm-bot: Finished rebooting node cloudcephosd2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 12:02 wm-bot: Rebooting node cloudcephosd2003-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 12:02 wm-bot: Finished rebooting node cloudcephosd2002-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:59 wm-bot: Rebooting node cloudcephosd2002-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:59 wm-bot: Finished rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:56 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:56 wm-bot: Rebooting the nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev - cookbook ran by dcaro@vulcanus * 11:55 wm-bot: Finished rebooting the nodes ['cloudcephmon2004-dev', 'cloudcephmon2005-dev', 'cloudcephmon2006-dev'] - cookbook ran by dcaro@vulcanus * 11:55 wm-bot: Finished rebooting node cloudcephmon2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:52 wm-bot: Rebooting node cloudcephmon2006-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:52 wm-bot: Finished rebooting node cloudcephmon2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:47 wm-bot: Rebooting node cloudcephmon2005-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:47 wm-bot: Finished rebooting node cloudcephmon2004-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:43 wm-bot: Rebooting node cloudcephmon2004-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 11:43 wm-bot: Rebooting the nodes cloudcephmon2004-dev,cloudcephmon2005-dev,cloudcephmon2006-dev - cookbook ran by dcaro@vulcanus === 2022-04-26 === * 10:36 taavi: [codfw1dev] updated designate pool to 2004/2005-dev according to the instructions on https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/DNS/Designate#Initial_designate/pdns_node_setup === 2022-04-22 === * 10:33 taavi: [codfw1dev] restart designate-sink on both new cloudservices host to fix rabbitmq connectivity === 2022-04-21 === * 05:38 andrewbogott: replaced cloudservices200[2,3] with cloudservices200[4,5] === 2022-04-19 === * 15:29 andrewbogott: stopping all VMs on cloudvirt1019, reimaging host === 2022-04-18 === * 15:23 andrewbogott: reimaging cloudvirt1020, leaving VMs in place * 13:40 andrewbogott: shutting down many codfdfw1dev servers (including network infra!) for [[phab:T305469|T305469]] === 2022-04-14 === * 20:14 andrewbogott: restarting nova-api and nova-conductor services in a superstitious attempt to reduce open DB connections === 2022-04-13 === * 22:01 andrewbogott: restarting galera on cloudcontrols (one by one) to clear open connections === 2022-04-11 === * 15:59 taavi: created cloudinfra.wmcloud.org zone === 2022-04-09 === * 19:55 andrewbogott: reimaging cloudbackup1001-dev to bullseye * 19:37 taavi: add 'puppet-enc' service & endpoint to keystone [[phab:T274666|T274666]] * 19:25 andrewbogott: reimaging cloudbackup1002-dev to bullseye === 2022-04-07 === * 12:51 wm-bot: Set cloudvirt 'cloudvirt1016.eqiad.wmnet' maintenance. ([[phab:T305631|T305631]]) - cookbook ran by arturo@nostromo === 2022-04-06 === * 09:12 arturo: [codf1dev] installing python3-eventlet 0.30.2-5~bpo11+1 on all required servers (cloudvirt, cloudnet, cloudcontrol) ([[phab:T305157|T305157]]) * 08:45 arturo: [codfw1dev] trying with python3-eventlet 0.30.2-5 installed by hand on cloudvirt2003-dev ([[phab:T305157|T305157]]) * 08:42 arturo: [codfw1dev] trying with python3-eventlet 0.30.2-5 installed by hand on cloudcontrol servers ([[phab:T305157|T305157]]) * 08:24 arturo: [codfw1dev] trying with python3-dnspython 2.2.0-2 installed by hand on cloudvirt2003-dev ([[phab:T305157|T305157]]) * 08:20 arturo: [codfw1dev] trying with python3-dnspython 2.2.0-2 installed by hand on cloudcontrol servers ([[phab:T305157|T305157]]) === 2022-03-30 === * 11:20 arturo: apply urpf strict filter to eqiad cloud-hosts vlan - [[phab:T285461|T285461]] === 2022-03-29 === * 10:02 dcaro: restarting keystone ([[phab:T304918|T304918]]) === 2022-03-23 === * 22:53 wm-bot: Drained 'cloudvirt1045.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:38 wm-bot: Drained 'cloudvirt1044.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:12 wm-bot: Set cloudvirt 'cloudvirt1045.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:12 wm-bot: Draining 'cloudvirt1045.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:08 wm-bot: Set cloudvirt 'cloudvirt1043.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:07 wm-bot: Set cloudvirt 'cloudvirt1044.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:06 wm-bot: Draining 'cloudvirt1044.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:06 wm-bot: Draining 'cloudvirt1043.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:54 wm-bot: Drained 'cloudvirt1042.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:19 wm-bot: Set cloudvirt 'cloudvirt1042.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:19 wm-bot: Draining 'cloudvirt1042.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:12 wm-bot: Drained 'cloudvirt1040.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:12 wm-bot: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:09 wm-bot: Draining 'cloudvirt1040.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:07 wm-bot: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 21:04 wm-bot: Draining 'cloudvirt1040.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:55 wm-bot: Set cloudvirt 'cloudvirt1041.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:54 wm-bot: Draining 'cloudvirt1041.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:30 wm-bot: Drained 'cloudvirt1039.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:15 wm-bot: Set cloudvirt 'cloudvirt1040.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:15 wm-bot: Set cloudvirt 'cloudvirt1039.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:14 wm-bot: Draining 'cloudvirt1040.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:14 wm-bot: Draining 'cloudvirt1039.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:44 wm-bot: Set cloudvirt 'cloudvirt1038.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:43 wm-bot: Draining 'cloudvirt1038.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:19 wm-bot: Set cloudvirt 'cloudvirt1037.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:18 wm-bot: Draining 'cloudvirt1037.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:13 wm-bot: Drained 'cloudvirt1036.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:02 wm-bot2: Testing wm-bot relay to #wikimedia-cloud-feed * 17:55 wm-bot: Set cloudvirt 'cloudvirt1036.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:54 wm-bot: Draining 'cloudvirt1036.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:04 wm-bot: Set cloudvirt 'cloudvirt1035.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:03 wm-bot: Draining 'cloudvirt1035.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:03 wm-bot: Drained 'cloudvirt1034.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:51 wm-bot: Drained 'cloudvirt1033.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:37 wm-bot: Set cloudvirt 'cloudvirt1034.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:37 wm-bot: Set cloudvirt 'cloudvirt1033.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:36 wm-bot: Draining 'cloudvirt1034.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:36 wm-bot: Draining 'cloudvirt1033.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 15:01 wm-bot: Drained 'cloudvirt1032.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 15:00 wm-bot: Set cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:57 wm-bot: Draining 'cloudvirt1032.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:44 wm-bot: Drained 'cloudvirt1031.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:35 wm-bot: Set cloudvirt 'cloudvirt1032.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:34 wm-bot: Draining 'cloudvirt1032.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:32 wm-bot: Drained 'cloudvirt1030.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:20 wm-bot: Set cloudvirt 'cloudvirt1031.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:19 wm-bot: Draining 'cloudvirt1031.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:18 wm-bot: Set cloudvirt 'cloudvirt1030.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 14:17 wm-bot: Draining 'cloudvirt1030.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 13:54 taavi: restart nova-fullstack on cloudcontrol1003 to pick up bastion ip change * 13:43 wm-bot: Drained 'cloudvirt1029.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 13:23 wm-bot: Set cloudvirt 'cloudvirt1029.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 13:22 wm-bot: Draining 'cloudvirt1029.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster === 2022-03-22 === * 22:59 wm-bot: Set cloudvirt 'cloudvirt1027.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 22:58 wm-bot: Draining 'cloudvirt1027.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster === 2022-03-17 === * 01:09 wm-bot: Drained 'cloudvirt1016.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 00:53 wm-bot: Set cloudvirt 'cloudvirt1016.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 00:52 wm-bot: Setting cloudvirt 'cloudvirt1016.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 00:52 wm-bot: Draining 'cloudvirt1016.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster === 2022-03-15 === * 20:58 wm-bot: Drained 'cloudvirt1026.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:36 wm-bot: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:36 wm-bot: Setting cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:36 wm-bot: Draining 'cloudvirt1026.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 13:14 wm-bot: Unset cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by arturo@nostromo * 13:14 wm-bot: Unsetting cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by arturo@nostromo * 10:32 wm-bot: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by arturo@nostromo * 10:30 wm-bot: Setting cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. - cookbook ran by arturo@nostromo === 2022-03-14 === * 21:24 wm-bot: Drained 'cloudvirt1025.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:59 wm-bot: Set cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:58 wm-bot: Setting cloudvirt 'cloudvirt1025.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:58 wm-bot: Draining 'cloudvirt1025.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:15 wm-bot: Setting cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:15 wm-bot: Draining 'cloudvirt1024.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 20:02 wm-bot: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 19:59 wm-bot: Setting cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 19:59 wm-bot: Draining 'cloudvirt1024.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 19:16 wm-bot: Set cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 19:15 wm-bot: Setting cloudvirt 'cloudvirt1024.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 19:15 wm-bot: Draining 'cloudvirt1024.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 19:13 wm-bot: Drained 'cloudvirt1023.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:56 wm-bot: Set cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:55 wm-bot: Setting cloudvirt 'cloudvirt1023.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:55 wm-bot: Draining 'cloudvirt1023.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:53 wm-bot: Drained 'cloudvirt1022.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:52 wm-bot: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:51 wm-bot: Setting cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:51 wm-bot: Draining 'cloudvirt1022.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:50 wm-bot: Drained 'cloudvirt1021.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:48 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:48 wm-bot: Setting cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:48 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 11:48 dcaro: rebased cookbooks on latest master, make sure you pull before sending new patches === 2022-03-08 === * 18:29 wm-bot: Set cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:29 wm-bot: Setting cloudvirt 'cloudvirt1022.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:29 wm-bot: Draining 'cloudvirt1022.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:23 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:21 wm-bot: Setting cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:21 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:18 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:17 wm-bot: Setting cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 18:17 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:28 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:27 wm-bot: Setting cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:27 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:18 wm-bot: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:15 wm-bot: Setting cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 17:15 wm-bot: Draining 'cloudvirt1017.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:48 wm-bot: Set cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:47 wm-bot: Setting cloudvirt 'cloudvirt1017.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:47 wm-bot: Draining 'cloudvirt1017.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:36 wm-bot: Drained 'cloudvirt1016.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:08 wm-bot: Set cloudvirt 'cloudvirt1016.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:07 wm-bot: Setting cloudvirt 'cloudvirt1016.eqiad.wmnet' maintenance. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 16:07 wm-bot: Draining 'cloudvirt1016.eqiad.wmnet'. ([[phab:T281276|T281276]]) - cookbook ran by andrew@buster * 13:11 arturo: [codfw1dev] rebooting cloudservices servers for [[phab:T303179|T303179]] * 13:07 arturo: [codfw1dev] rebooting cloudvirt servers for [[phab:T303179|T303179]] * 13:06 arturo: [codfw1dev] rebooting cloudnet servers for [[phab:T303179|T303179]] * 12:55 arturo: [codfw1dev] rebooting cloudcontrol servers for [[phab:T303179|T303179]] === 2022-03-03 === * 08:49 taavi: deploying cloudmetrics grafana to grafana 8, [[phab:T282863|T282863]] === 2022-03-02 === * 09:06 arturo: merging core router firewall change https://gerrit.wikimedia.org/r/c/operations/homer/public/+/701347 === 2022-02-28 === * 15:30 dcaro: cleaning up leftover snapshots from failed backups of the maps volume ([[phab:T302720|T302720]]) === 2022-02-24 === * 17:04 andrewbogott: upgrading eqiad1 and codfw1dev to mariadb 10.5.15+maria~bullseye via 'apt-get install libmariadb3:amd64 galera-4 mariadb-server' * 15:42 dcaro: stopping and starting mariadb on cloudcontrol1003 ([[phab:T302146|T302146]]) * 10:37 arturo: [codfw1dev] briefly installed galera-4 (26.4.11+1bullseye) over (26.4.9-0+deb11u1) on cloudcontrol2001-dev and then downgrade again to verify package install ([[phab:T302482|T302482]]) === 2022-02-23 === * 20:39 taavi: added domain-wide 'designateadmin' and 'observer' roles to project-proxy-dns-manager service account [[phab:T295246|T295246]] * 17:40 andrewbogott: restarting lots of openstack services to try to clear up the mess that is [[phab:T236101|T236101]] * 12:13 arturo: cleaning up cinder volume snapshots, aborrero@cloudcontrol1005:~$ for i in $(sudo wmcs-openstack volume snapshot list -f value -c ID) ; do sudo wmcs-openstack volume snapshot delete $i ; done ([[phab:T302382|T302382]]) * 10:14 arturo: cleaning up neutron agents for non-existent servers cloudvirt100[1-9].eqiad.wmnet,cloudvirt10[12-15].eqiad.wmnet * 10:05 dcaro: Deleting stuck novafullstack servers, to let the service create new ones ([[phab:T302369|T302369]]) * 09:56 arturo: neutron agent-delete bad663b3-fd25-4393-a546-{{Gerrit|4b1b4bdec4db}} (Linux bridge agent {{!}} cloudvirtan1001) * 09:56 arturo: neutron agent-delete 1071c198-ed57-4b5a-9439-{{Gerrit|30e66a31aa69}} (Linux bridge agent {{!}} cloudvirtan1005) * 09:55 arturo: neutron agent-delete 2eeef198-8af7-4e5d-bd73-{{Gerrit|e14a2a8d2404}} (Linux bridge agent {{!}} cloudvirtan1004) * 09:55 arturo: neutron agent-delete afe173eb-35ba-444a-9960-{{Gerrit|899629786d2f}} (Linux bridge agent {{!}} cloudvirtan1003) * 09:54 arturo: neutron agent-delete afcb9b7f-c1a6-4ff4-9b10-{{Gerrit|92bfbe8d1a56}} (Linux bridge agent {{!}} cloudvirtan1002) * 09:39 dcaro: restarting neutron-api cloudcontrol1003 to see if the agent status update starts working ([[phab:T302369|T302369]]) * 09:38 dcaro: restarting neutron-dhcp-agent on cloudnet1003 ([[phab:T302369|T302369]]) === 2022-02-22 === * 22:10 andrewbogott: raising project 'maps' quota by two tb -- [[phab:T300160|T300160]] * 09:24 arturo: restarting mariadb @ cloudcontrol1003 ([[phab:T302146|T302146]]) * 09:13 arturo: restarting mariadb @ cloudcontrol1004 ([[phab:T302146|T302146]]) === 2022-02-18 === * 21:57 andrewbogott: leaving cloudcontrol1003 downtimed with disabled puppet for the weekend. Everything there should be stable and fine save rabbit which needs an upgrade. * 21:30 andrewbogott: rebooting cloudcontrol1003 because rabbit is freaking out * 17:25 andrewbogott: in-place upgrade of cloudcontrol1004 to bullseye -- [[phab:T281276|T281276]] * 12:34 arturo: manually install prometheus-openstack-exporter on cloudcontrol1005 ([[phab:T302050|T302050]]) === 2022-02-17 === * 23:02 andrewbogott: in-place upgrade to Bullseye on cloudcontrol1005 [[phab:T281276|T281276]] === 2022-02-15 === * 14:15 taavi: [codfw1dev] added domain-wide 'designateadmin' and 'observer' roles to codfw1dev-proxy-dns-manager service account [[phab:T295246|T295246]] === 2022-02-04 === * 10:12 arturo: restart backup_vms service in cloudvirt1024 ([[phab:T300956|T300956]]) === 2022-02-03 === * 08:21 taavi: cloudmetrics1004: manually added an empty line to /etc/prometheus/blackbox.yml to make /usr/local/bin/blackbox-exporter-assemble happy (clearing "performing a change every puppet run" alert) === 2022-02-02 === * 02:36 andrewbogott: restarting mariadb on cloudcontrol1004 === 2022-01-31 === * 10:15 arturo: cloudcontrol1005:~$ sudo systemctl restart backup_glance_images.service (failed state, no logs, icinga alert) === 2022-01-29 === * 18:24 taavi: delete 2 puppet prefixes in a weird state [[phab:T299750|T299750]] === 2022-01-27 === * 13:24 arturo: cloudmetrics1004:~ $ sudo systemctl restart wmcs_monitoring_graphite_rsync.service ([[phab:T300138|T300138]]) === 2022-01-26 === * 19:09 andrewbogott: bootstrapping a fresh galera node on cloudcontrol1004 * 18:57 andrewbogott: restarting mariadb on cloudcontrol1004 === 2022-01-25 === * 10:49 arturo: made cloudmetrics1001/1002 primary/backup respectively ([[phab:T299744|T299744]], [[phab:T297814|T297814]], [[phab:T300011|T300011]]) === 2022-01-19 === * 16:38 andrewbogott: moving all scratch mounts to scratch.svc.cloudinfra-nfs.eqiad1.wikimedia.cloud === 2022-01-05 === * 03:11 andrewbogott: 'cp /etc/apt/sources.list /etc/apt/sources.list.prepuppet' on all VMs. Backing up state before puppetizing sources.list with https://gerrit.wikimedia.org/r/c/operations/puppet/+/751498 === 2022-01-04 === * 12:44 dcaro: increasing the size_limit for labs ldap servers === 2021-12-26 === * 16:55 majavah: run attachLdapUser.php on wikitech for developer account "Karthiksripal" === 2021-12-24 === * 22:51 majavah: ran the wikireplica dns script on s5 [[phab:T298303|T298303]] === 2021-12-23 === * 21:42 majavah: deployed horizon wmf-proxy-dashboard update to fix editing of existing proxies === 2021-12-21 === * 10:39 arturo: dropped egress NAT exceptions for WMF apt repos, [[phab:T298042|T298042]] === 2021-12-15 === * 12:44 dcaro: Downtiming cloudvirt-wdqs1001 as it has no VMs running until disk space is fixed ([[phab:T297454|T297454]]) === 2021-12-14 === * 10:26 dcaro: Moved the nova cache (/var/lib/nova/instances/_base) and the canary image local data (/var/lib/nova/instance/<canary_image_id>) to the root disk on cloudvirt-wdqs1001 to temporary free some space ([[phab:T297454|T297454]]) === 2021-12-13 === * 18:08 wm-bot: Drained 'cloudvirt1014.eqiad.wmnet'. - cookbook ran by michael@mouse * 17:50 wm-bot: Set cloudvirt 'cloudvirt1014.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse * 17:49 wm-bot: Setting cloudvirt 'cloudvirt1014.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse * 17:49 wm-bot: Draining 'cloudvirt1014.eqiad.wmnet'. - cookbook ran by michael@mouse * 17:44 wm-bot: Drained 'cloudvirt1013.eqiad.wmnet'. - cookbook ran by michael@mouse * 17:30 wm-bot: Set cloudvirt 'cloudvirt1013.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse * 17:30 wm-bot: Setting cloudvirt 'cloudvirt1013.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse * 17:30 wm-bot: Draining 'cloudvirt1013.eqiad.wmnet'. - cookbook ran by michael@mouse * 17:13 wm-bot: Drained 'cloudvirt1012.eqiad.wmnet'. - cookbook ran by michael@mouse * 16:50 wm-bot: Set cloudvirt 'cloudvirt1012.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse * 16:47 wm-bot: Setting cloudvirt 'cloudvirt1012.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse * 16:47 wm-bot: Draining 'cloudvirt1012.eqiad.wmnet'. - cookbook ran by michael@mouse * 16:44 wm-bot: Set cloudvirt 'cloudvirt1012.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse * 16:43 wm-bot: Setting cloudvirt 'cloudvirt1012.eqiad.wmnet' maintenance. - cookbook ran by michael@mouse === 2021-12-03 === * 18:56 andrewbogott: maintain-views and maintain-meta-p on clouddb1013-1020 * 10:49 majavah: deleting dbbackups-dashboard project [[phab:T296992|T296992]] === 2021-12-02 === * 01:17 wm-bot: Drained 'cloudvirt1028.eqiad.wmnet'. ([[phab:T296790|T296790]]) - cookbook ran by andrew@buster * 00:56 wm-bot: Set cloudvirt 'cloudvirt1028.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:56 wm-bot: Setting cloudvirt 'cloudvirt1028.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:56 wm-bot: Draining 'cloudvirt1028.eqiad.wmnet'. ([[phab:T296790|T296790]]) - cookbook ran by andrew@buster * 00:50 wm-bot: Set cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:50 wm-bot: Setting cloudvirt 'cloudvirt1026.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:50 wm-bot: Draining 'cloudvirt1026.eqiad.wmnet'. ([[phab:T296790|T296790]]) - cookbook ran by andrew@buster * 00:28 wm-bot: Drained 'cloudvirt1021.eqiad.wmnet'. ([[phab:T296790|T296790]]) - cookbook ran by andrew@buster * 00:03 wm-bot: Set cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:02 wm-bot: Setting cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 00:02 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T296790|T296790]]) - cookbook ran by andrew@buster === 2021-12-01 === * 23:59 wm-bot: Setting cloudvirt 'cloudvirt1021.eqiad.wmnet' maintenance. - cookbook ran by andrew@buster * 23:59 wm-bot: Draining 'cloudvirt1021.eqiad.wmnet'. ([[phab:T296790|T296790]]) - cookbook ran by andrew@buster * 23:54 andrewbogott: *correction* adding spare cloudvirts 1044 and 1045 to the 'ceph' pool in order to make space for future juggling around [[phab:T296790|T296790]] and [[phab:T296792|T296792]] * 23:53 andrewbogott: adding spare cloudvirts 1044 and 1055 to the 'ceph' pool in order to make space for future juggling around [[phab:T296790|T296790]] and [[phab:T296792|T296792]] === 2021-11-28 === * 17:48 andrewbogott: moved cloudvirt1018 out of the 'localstorage' aggregate and into 'maintenance' for [[phab:T296592|T296592]]. It will need to be moved back after the raid is rebuilt. === 2021-11-21 === * 07:19 dcaro_away: restarting designate-sink with some extra logs in it ([[phab:T296144|T296144]]) === 2021-11-17 === * 15:48 andrewbogott: upgrading mariadb packages on eqiad1 cloudcontrols * 15:39 andrewbogott: sudo cumin "cloud*" 'apt-get update -y --allow-releaseinfo-change' * 15:26 andrewbogott: updated mariadb packages on codfw1dev cloudcontrols to 1:10.3.31-0+deb10u1 === 2021-11-12 === * 13:31 arturo: restarting glance-api services to make sure they work with new ceph auth creds ([[phab:T293752|T293752]]) === 2021-11-08 === * 21:50 andrewbogott: returned clouddb pools back to normal after maintain_views run: https://gerrit.wikimedia.org/r/c/operations/puppet/+/737505 [[phab:T216481|T216481]] * 20:07 andrewbogott: depooling clouddb1013 for maintain_views attempt * 10:54 arturo: [codfw1dev] create service account `srv-networktests` following https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Service_accounts for [[phab:T294955|T294955]] * 10:34 arturo: create service account `srv-networktests` following https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Service_accounts for [[phab:T294955|T294955]] === 2021-11-05 === * 11:18 wm-bot: Added 1 new OSDs ['cloudcephosd1024.eqiad.wmnet'] ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:17 wm-bot: Added OSD cloudcephosd1024.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:15 wm-bot: Finished rebooting node cloudcephosd1024.eqiad.wmnet - cookbook ran by arturo@endurance * 11:12 wm-bot: Rebooting node cloudcephosd1024.eqiad.wmnet - cookbook ran by arturo@endurance * 11:12 wm-bot: Adding OSD cloudcephosd1024.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:12 wm-bot: Adding new OSDs ['cloudcephosd1024.eqiad.wmnet'] to the cluster ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance === 2021-11-04 === * 16:39 wm-bot: Added 1 new OSDs ['cloudcephosd1023.eqiad.wmnet'] ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:39 wm-bot: Added OSD cloudcephosd1023.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:37 wm-bot: Finished rebooting node cloudcephosd1023.eqiad.wmnet - cookbook ran by arturo@endurance * 16:34 wm-bot: Rebooting node cloudcephosd1023.eqiad.wmnet - cookbook ran by arturo@endurance * 16:33 wm-bot: Adding OSD cloudcephosd1023.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:33 wm-bot: Adding new OSDs ['cloudcephosd1023.eqiad.wmnet'] to the cluster ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:17 wm-bot: Added 1 new OSDs ['cloudcephosd1022.eqiad.wmnet'] ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:17 wm-bot: Added OSD cloudcephosd1022.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:16 wm-bot: Finished rebooting node cloudcephosd1022.eqiad.wmnet - cookbook ran by arturo@endurance * 16:13 wm-bot: Rebooting node cloudcephosd1022.eqiad.wmnet - cookbook ran by arturo@endurance * 16:12 wm-bot: Adding OSD cloudcephosd1022.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:12 wm-bot: Adding new OSDs ['cloudcephosd1022.eqiad.wmnet'] to the cluster ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:00 wm-bot: Adding OSD cloudcephosd1022.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 16:00 wm-bot: Adding new OSDs ['cloudcephosd1022.eqiad.wmnet'] to the cluster ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:26 wm-bot: Added 1 new OSDs ['cloudcephosd1021.eqiad.wmnet'] ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:26 wm-bot: Added OSD cloudcephosd1021.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:23 wm-bot: Finished rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by arturo@endurance * 11:20 wm-bot: Rebooting node cloudcephosd1021.eqiad.wmnet - cookbook ran by arturo@endurance * 11:19 wm-bot: Adding OSD cloudcephosd1021.eqiad.wmnet... (1/1) ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:19 wm-bot: Adding new OSDs ['cloudcephosd1021.eqiad.wmnet'] to the cluster ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance * 11:16 wm-bot: Adding new OSDs ['cloudcephosd1021.eqiad.wmnet'] to the cluster ([[phab:T295012|T295012]]) - cookbook ran by arturo@endurance === 2021-11-03 === * 17:22 arturo: [codfw1dev] installing keepalived 2.1.5 from buster-backports on cloudgw2001-dev/2002-dev ([[phab:T294956|T294956]]) * 11:45 arturo: [codfw1dev] downgrade kernel on cloudgw2001-dev/2002-dev ([[phab:T294853|T294853]], [[phab:T291813|T291813]]) === 2021-11-02 === * 10:54 arturo: rebooting cloudnet1004/1003 for [[phab:T291813|T291813]] * 10:43 arturo: [codfw1dev] rebooting cloudgw200[12]-dev for [[phab:T291813|T291813]] === 2021-10-24 === * 00:47 andrewbogott: deploying a change so that openstack clients use tls endpoints: https://gerrit.wikimedia.org/r/c/operations/puppet/+/732738 === 2021-10-21 === * 10:19 arturo: drop firewall exception on core routers for wiki replicas legacy setup ([[phab:T293897|T293897]]) * 10:12 arturo: drop NAT exception for wiki replicas legacy setup ([[phab:T293897|T293897]]) === 2021-10-20 === * 21:06 andrewbogott: creating cloudinfra-nfs project [[phab:T293936|T293936]] === 2021-10-18 === * 19:21 andrewbogott: also ticked the 'admin' box on wikitech for majavah [[phab:T292827|T292827]] * 18:58 andrewbogott: granting majavah 'admin' role in the 'admin' project and also in the default domain. [[phab:T292827|T292827]] === 2021-10-14 === * 12:28 arturo: [codfw1dev] add DB grants for cloudbackup2002.codfw.wmnet IP address to the cinder DB ([[phab:T292546|T292546]]) === 2021-10-13 === * 10:46 arturo: updating python3-neutron across the fleet ([[phab:T292936|T292936]]) === 2021-10-12 === * 09:06 dcaro: upgrading eqiad cloudnet hosts neutron packages ([[phab:T292936|T292936]]) * 08:57 dcaro: upgrading codfw cloudnet hosts neutron packages ([[phab:T292936|T292936]]) === 2021-10-05 === * 09:39 arturo: [codfw1dev] cleaning up manila stuff from openstack (db, endpoints, tenant, VMs, and such) [[phab:T291257|T291257]] === 2021-09-30 === * 14:50 andrewbogott: sudo cumin "cloud*" "ps -ef {{!}} grep nslcd && service nslcd restart" and sudo cumin "lab*" "ps -ef {{!}} grep nslcd && service nslcd restart" [[phab:T292202|T292202]] * 14:43 andrewbogott: ran sudo cumin --force --timeout 500 -o json "A:all" "ps -ef {{!}} grep nslcd && service nslcd restart" to get nslcd happy again [[phab:T292202|T292202]] === 2021-09-29 === * 09:41 arturo: [codfw1dev] cleanup manila shares definitions for a clean start now that the manila-sharecontroller VM is apparently well configured ([[phab:T291257|T291257]]) === 2021-09-28 === * 16:23 bstorm: downtime for clouddb1020 to reduce re-pages in case this goes badly [[phab:T291963|T291963]] * 16:21 bstorm: powering on clouddb1020 via remote console [[phab:T291963|T291963]] * 15:58 bstorm: depooled clouddb1020 for repair [[phab:T291961|T291961]] * 12:40 dcaro: Merged change on sssd for bullseye cloud hosts ([[phab:T291585|T291585]]) * 11:30 arturo: [codfw1dev] create floating IP 185.15.57.5 for manila-sharecontroller.cloudinfra-codfw1dev.codfw1dev.wmcloud.org ([[phab:T291257|T291257]]) === 2021-09-27 === * 10:07 arturo: cloudcontrol1004 apparently healthy [[phab:T291446|T291446]] * 09:25 arturo: rebooting cloudcontrol1004 for [[phab:T291446|T291446]] === 2021-09-24 === * 13:02 arturo: [codfw1dev] create VM manila-share-controller-01 on cloudinfra-codfw1dev * 13:00 arturo: [codfw1dev] rebase labs/private.git on cloudinfra-puppetmaster-01, had merge conflict === 2021-09-21 === * 12:13 arturo: [codfw1dev] trying to create a manila service image ([[phab:T291257|T291257]]) * 11:45 arturo: [codfw1dev] created rabbitmq user ([[phab:T291257|T291257]]) * 11:32 arturo: [codfw1dev] populated manila DB & created service endpoints ([[phab:T291257|T291257]]) * 11:06 arturo: [codfw1dev] give manila user admin role @ manila project ([[phab:T291257|T291257]]) * 11:06 arturo: [codfw1dev] created manila project ([[phab:T291257|T291257]]) * 10:57 arturo: [codfw1dev] created manila user @ labtestwikitech ([[phab:T291257|T291257]]) * 10:49 arturo: [codfw1dev] create manila database on cloudcontrol-dev nodes (galera) [[phab:T291257|T291257]] === 2021-09-20 === * 23:08 bstorm: ran `echo check > /sys/block/md0/md/sync_action` on cloudcontrol1004 to check raid * 22:48 andrewbogott: stopped puppet & mariadb on cloudcontrol1004; it was flapping * 22:44 andrewbogott: sudo touch /tmp/galera.disabled on cloudcontrol1004, the service seems troubled there * 21:57 andrewbogott: moving cloudvirt1043 into the 'nfs' aggregate for [[phab:T291405|T291405]] === 2021-09-17 === * 11:35 arturo: [codfw1dev] install manila on cloudcontrol2001-dev ([[phab:T291257|T291257]]) === 2021-09-16 === * 15:56 bstorm: removing downtime for labstore1005 so we'll know if it has another issue [[phab:T290318|T290318]] === 2021-09-09 === * 22:03 bstorm: restarted the prometheus-mysqld-exporter@s1 service as it was not working [[phab:T290630|T290630]] * 03:15 bstorm: resetting swap on clouddb1017 [[phab:T290630|T290630]] * 03:08 andrewbogott: stopping maintain-dbusers on labstore1004 for help diagnosing [[phab:T290630|T290630]] === 2021-09-03 === * 15:34 bstorm: rebooting labstore1005 to disconnect the drives from labstore1004 [[phab:T290318|T290318]] * 15:24 bstorm: stopping puppet and disabling backup syncs to labstore1005 on cloudbackup2002 [[phab:T290318|T290318]] * 15:20 bstorm: stopping puppet and disabling backup syncs to labstore1005 on cloudbackup2001 [[phab:T290318|T290318]] === 2021-08-30 === * 16:16 wm-bot: Added 1 new OSDs ['cloudcephosd1018.eqiad.wmnet'] - cookbook ran by andrew@buster * 16:16 wm-bot: Added OSD cloudcephosd1018.eqiad.wmnet... (1/1) - cookbook ran by andrew@buster * 16:13 wm-bot: Adding OSD cloudcephosd1018.eqiad.wmnet... (1/1) - cookbook ran by andrew@buster * 16:13 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster * 16:10 wm-bot: Finished rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by andrew@buster * 16:07 wm-bot: Rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by andrew@buster * 16:07 wm-bot: Adding OSD cloudcephosd1018.eqiad.wmnet... (1/1) - cookbook ran by andrew@buster * 16:07 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster === 2021-08-27 === * 18:57 andrewbogott: raising toolsbeta ram/core/instances quotas so majavah can experiment with bullseye === 2021-08-25 === * 14:45 wm-bot: Finished rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by andrew@buster * 14:42 wm-bot: Rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by andrew@buster * 14:42 wm-bot: Adding OSD cloudcephosd1018.eqiad.wmnet... (1/1) - cookbook ran by andrew@buster * 14:42 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster * 14:41 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster === 2021-08-19 === * 17:39 bstorm: restarting glance image backup to try and clear the page === 2021-08-18 === * 16:21 wm-bot: Rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by andrew@buster * 16:21 wm-bot: Adding OSD cloudcephosd1018.eqiad.wmnet... (1/1) - cookbook ran by andrew@buster * 16:21 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster * 16:17 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster * 16:16 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster * 16:15 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster * 16:13 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster - cookbook ran by andrew@buster * 14:47 andrewbogott: adding clouvirt1038 to the ceph aggregate, removing from the maintenance aggregate [[phab:T276922|T276922]] === 2021-08-17 === * 15:11 andrewbogott: rebooting cloudcephosd1008 to force raid rebuild -- [[phab:T287838|T287838]] === 2021-08-11 === * 13:51 wm-bot: Finished rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:48 wm-bot: Rebooting node cloudcephosd1018.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 13:47 wm-bot: Adding OSD cloudcephosd1018.eqiad.wmnet... (1/1) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 13:47 wm-bot: Adding new OSDs ['cloudcephosd1018.eqiad.wmnet'] to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus === 2021-08-10 === * 15:15 andrewbogott: restarting all designate services in eqiad1 * 15:04 andrewbogott: restarting designate-sink in eqiad1; it's complaining about rabbit but I don't want to restart rabbit yet === 2021-08-05 === * 09:37 dcaro: Taking one osd daemon down ot codfw cluster ([[phab:T288203|T288203]]) === 2021-08-04 === * 19:20 bd808: Running deleteBatch.php on cloudweb2001-dev to remove legacy Heira: pages from labtestwiki === 2021-08-03 === * 17:40 bstorm: rerunning the glance backup script after failure === 2021-07-31 === * 00:10 andrewbogott: "systemctl reset-failed cloud-init.service" on all VMs for [[phab:T287309|T287309]] * 00:08 andrewbogott: "systemctl reset-failed cloud-final.service" on all VMs for [[phab:T287309|T287309]] === 2021-07-27 === * 21:32 andrewbogott: putting cloudvirt1012 back into service [[phab:T286748|T286748]] * 20:52 andrewbogott: draining VMs off of cloudvirt1012 so we can replace the battery for [[phab:T286748|T286748]] * 15:15 andrewbogott: "rm /etc/apt/sources.list.d/openstack-mitaka-jessie.list" cloud-wide === 2021-07-23 === * 15:22 bstorm: update wikireplicas-dns for s7 fix for web replicas === 2021-07-20 === * 17:07 andrewbogott: reloading haproxy on dbproxy1018 for [[phab:T286598|T286598]] * 15:45 arturo: failback from labstore1006 to labstore1007 (dumps NFS) https://gerrit.wikimedia.org/r/c/operations/puppet/+/705417 * 00:10 bstorm: restarting nova-api on cloudcontrol1003 to try and recover whatever it's doing with designate_floating_ip_ptr_records_updater === 2021-07-19 === * 22:05 bstorm: set downtime scheduled for tomorrow from 1300 to 1600 UTC for cloudstore1008 and 1009 [[phab:T286599|T286599]] * 20:40 andrewbogott: reloading haproxy on dbproxy1018 for [[phab:T286598|T286598]] * 13:50 andrewbogott: upgrading mariadb to 10.3.29 on all cloudcontrols === 2021-07-16 === * 09:55 dcaro: checking HP raid issues on coludvirt1012 ([[phab:T286766|T286766]]) === 2021-07-14 === * 21:08 andrewbogott: restarting lots of openstack services while trying to resolve [[phab:T286675|T286675]] * 12:17 dcaro: doing ceph outage tests on codfw1 (fyi) === 2021-07-13 === * 10:57 dcaro: enabled autoscaling on codfw1 ceph cluster, setting a minimum of pgs on codfw1dev-compute to 128 === 2021-07-02 === * 10:12 wm-bot: The cluster is not rebalance after adding the new OSDs ['cloudcephosd1019.eqiad.wmnet', 'cloudcephosd1020.eqiad.wmnet'] ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:12 wm-bot: Added 2 new OSDs ['cloudcephosd1019.eqiad.wmnet', 'cloudcephosd1020.eqiad.wmnet'] ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:12 wm-bot: Added OSD cloudcephosd1020.eqiad.wmnet... (2/2) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:10 wm-bot: Finished rebooting node cloudcephosd1020.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 10:07 wm-bot: Rebooting node cloudcephosd1020.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 10:07 wm-bot: Adding OSD cloudcephosd1020.eqiad.wmnet... (2/2) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:07 wm-bot: Added OSD cloudcephosd1019.eqiad.wmnet... (1/2) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:05 wm-bot: Finished rebooting node cloudcephosd1019.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 10:02 wm-bot: Rebooting node cloudcephosd1019.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 10:02 wm-bot: Adding OSD cloudcephosd1019.eqiad.wmnet... (1/2) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:01 wm-bot: Adding new OSDs ['cloudcephosd1019.eqiad.wmnet', 'cloudcephosd1020.eqiad.wmnet'] to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 09:13 wm-bot: Adding OSD cloudcephosd1019.eqiad.wmnet... (1/2) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 09:13 wm-bot: Adding new OSDs ['cloudcephosd1019.eqiad.wmnet', 'cloudcephosd1020.eqiad.wmnet'] to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus === 2021-07-01 === * 16:27 bstorm: failed over cloudstore1009 to cloudstore1008 [[phab:T224747|T224747]] * 16:18 bstorm: downtimed cloudstore1008 and cloudstore1009 to fail over [[phab:T224747|T224747]] * 14:25 wm-bot: Adding OSD cloudcephosd1019.eqiad.wmnet... (2/3) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 14:25 wm-bot: Added OSD cloudcephosd1017.eqiad.wmnet... (1/3) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 14:24 wm-bot: Finished rebooting node cloudcephosd1017.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:21 wm-bot: Rebooting node cloudcephosd1017.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:20 wm-bot: Adding OSD cloudcephosd1017.eqiad.wmnet... (1/3) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 14:20 wm-bot: Adding new OSDs ['cloudcephosd1017.eqiad.wmnet', 'cloudcephosd1019.eqiad.wmnet', 'cloudcephosd1020.eqiad.wmnet'] to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 14:18 wm-bot: Rebooting node cloudcephosd1017.eqiad.wmnet - cookbook ran by dcaro@vulcanus * 14:17 wm-bot: Adding OSD cloudcephosd1017.eqiad.wmnet... (1/3) ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 14:17 wm-bot: Adding new OSDs ['cloudcephosd1017.eqiad.wmnet', 'cloudcephosd1019.eqiad.wmnet', 'cloudcephosd1020.eqiad.wmnet'] to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 11:16 wm-bot: Added new OSD node cloudcephosd1016.eqiad.wmnet ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 11:13 wm-bot: Adding new OSD cloudcephosd1016.eqiad.wmnet to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:58 dcaro: rebooting cloudcephosd1016 ([[phab:T285858|T285858]]) * 10:47 wm-bot: Adding new OSD cloudcephosd1016.eqiad.wmnet to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:44 wm-bot: Adding new OSD cloudcephosd1016.eqiad.wmnet to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:42 wm-bot: Adding new OSD cloudcephosd1016.eqiad.wmnet to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:41 wm-bot: Adding new OSD cloudcephosd1016.eqiad.wmnet to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus * 10:40 wm-bot: Adding new OSD cloudcephosd1016.eqiad.wmnet to the cluster ([[phab:T285858|T285858]]) - cookbook ran by dcaro@vulcanus === 2021-06-30 === * 21:48 bstorm: downtimed space alerts for scratch on cloudstore1008 until after the migration === 2021-06-25 === * 15:28 andrewbogott: restarting openstack services on cloudcontrol1005 * 09:16 arturo: icinga downtime cloudcontrols for 2h * 08:20 dcaro: restarting rabbitmq on cloudcontrol100<nowiki>{</nowiki>3,4<nowiki>}</nowiki> === 2021-06-21 === * 13:54 dcaro: puppet fix merged and deployed, servers are back to normal * 13:20 dcaro: merged broken puppet patch, downtimed all cloudvirts for 2h while fixing (nothing big, just added a bad systemd timer) === 2021-06-20 === * 22:21 andrewbogott: clearing admin-monitoring VMs; puppet has been failing lately due to a full drive on the puppetmaster === 2021-06-15 === * 01:18 bstorm: running a modified version of the prometheus dir size cron in screen [[phab:T284964|T284964]] === 2021-06-14 === * 10:13 dcaro: setting ssd to debug mode on tools-sgeexec-0917 ([[phab:T284130|T284130]]) === 2021-06-10 === * 10:58 wm-bot: Finished rebooting the nodes ['cloudcephmon2002-dev', 'cloudcephmon2003-dev', 'cloudcephmon2004-dev'] ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:58 wm-bot: Finished rebooting node cloudcephmon2004-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:55 wm-bot: Rebooting node cloudcephmon2004-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:55 wm-bot: Finished rebooting node cloudcephmon2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:52 wm-bot: Rebooting node cloudcephmon2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:52 wm-bot: Finished rebooting node cloudcephmon2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:49 wm-bot: Rebooting node cloudcephmon2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:49 wm-bot: Rebooting the nodes cloudcephmon2002-dev,cloudcephmon2003-dev,cloudcephmon2004-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:48 wm-bot: Finished rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:48 wm-bot: Finished rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:45 wm-bot: Rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:45 wm-bot: Finished rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:42 wm-bot: Rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:42 wm-bot: Finished rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:39 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 10:39 wm-bot: Rebooting the nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:39 wm-bot: Finished rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:38 wm-bot: Finished rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:35 wm-bot: Rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:35 wm-bot: Finished rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:32 wm-bot: Rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:32 wm-bot: Finished rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:29 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:29 wm-bot: Rebooting the nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:26 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:26 wm-bot: Rebooting the nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:24 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 09:24 wm-bot: Rebooting the nodes cloudcephosd2001-dev,cloudcephosd2002-dev,cloudcephosd2003-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus === 2021-06-09 === * 17:33 arturo: removed icinga downtime for cloudmetrics1002 -- to see if hardware is healthy ([[phab:T281881|T281881]]) * 13:30 wm-bot: Finished rebooting the nodes ['cloudcephmon2002-dev', 'cloudcephmon2003-dev', 'cloudcephmon2004-dev'] ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:30 wm-bot: Finished rebooting node cloudcephmon2004-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:27 wm-bot: Rebooting node cloudcephmon2004-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:27 wm-bot: Finished rebooting node cloudcephmon2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:24 wm-bot: Rebooting node cloudcephmon2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:24 wm-bot: Finished rebooting node cloudcephmon2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:21 wm-bot: Rebooting node cloudcephmon2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:21 wm-bot: Rebooting the nodes cloudcephmon2002-dev,cloudcephmon2003-dev,cloudcephmon2004-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:01 wm-bot: Rebooting node cloudcephmon2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 13:01 wm-bot: Rebooting the nodes cloudcephmon2002-dev,cloudcephmon2003-dev,cloudcephmon2004-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 12:53 wm-bot: Rebooting node cloudcephmon2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 12:53 wm-bot: Rebooting the nodes cloudcephmon2002-dev,cloudcephmon2003-dev,cloudcephmon2004-dev ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus === 2021-06-08 === * 23:19 bd808: Downtimed cloudmetrics1002 in icinga until 2021-06-30 23:59:01 ([[phab:T281881|T281881]]) * 21:08 bstorm: downtiming grafana-labs for maintenance * 16:28 wm-bot: Finished rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:27 wm-bot: Finished rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:24 wm-bot: Rebooting node cloudcephosd2003-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:24 wm-bot: Finished rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:22 wm-bot: Rebooting node cloudcephosd2002-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:21 wm-bot: Finished rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:18 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:18 wm-bot: Rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 16:17 wm-bot: Rebooting the nodes ['cloudcephosd2001-dev', 'cloudcephosd2002-dev', 'cloudcephosd2003-dev'] ([[phab:T281248|T281248]]) - cookbook ran by dcaro@vulcanus * 15:03 wm-bot: Finished rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:59 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:59 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:57 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:57 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:29 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:23 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus * 14:18 wm-bot: Rebooting node cloudcephosd2001-dev.codfw.wmnet - cookbook ran by dcaro@vulcanus === 2021-06-07 === * 14:27 andrewbogott: moving cloudvirt1040 from 'maintenance' aggregate to 'ceph' aggregate [[phab:T281399|T281399]] === 2021-06-01 === * 13:12 dcaro: Changed the ceph osd_memory_target on eqiad pool to 6Gi (we were reaching the limit, swapping at some points) * 09:57 arturo: fix PTR record for 185.15.56.1 ([[phab:T284025|T284025]]) * 09:56 arturo: fix PTR record for 185.15.56.1 ([[phab:T248025|T248025]]) === 2021-05-27 === * 14:58 wm-bot: Testing - cookbook ran by dcaro@vulcanus === 2021-05-26 === * 19:10 andrewbogott: reimaging cloudvirt1018 to support local VM storage * 18:07 andrewbogott: draining cloudvirt1018, converting it to a local-storage host like cloudvirt1019 and 1020 -- [[phab:T283296|T283296]] * 14:36 dcaro: Enabled syslog logging for osd.55 on eqiad ceph cluster for testing ([[phab:T281247|T281247]]) * 14:36 dcaro: Enabled syslog logging on codfw ceph cluster (mon/osd/mgr) ([[phab:T281247|T281247]]) * 11:26 arturo: [codfw1dev] purge old kernel packages in cloudvirt200[12]-dev * 11:03 arturo: created public flavor `g3.cores16.ram36.disk20` (even though it was requested as private in [[phab:T283293|T283293]], but may be useful for others) === 2021-05-25 === * 16:14 bd808: Closed #wikimedia-cloud-admin on f***node * 16:11 bd808: Closed #wikimedia-cloud-feed on f***node * 15:19 dcaro: rebooted cloudvirt1020, starting VMs ([[phab:T275893|T275893]]) * 15:13 dcaro: rebooting cloudvirt1020 ([[phab:T275893|T275893]]) * 14:42 dcaro: taking cloudvirt1020 out for maintenance (openstack wise) so no new VMs are scheduled on it ([[phab:T275893|T275893]]) === 2021-05-24 === * 22:32 andrewbogott: changing the default ttl for eqiad1.wikimedia.cloud. from 3600 to 60; this should help us avoid madness when re-using hostnames. * 11:20 arturo: created `g3.cores2.ram80.disk40.private` for the wmf-research-tools project, to allow resizing a 40G disk instance === 2021-05-22 === * 02:14 bstorm: downtiming SMART alerts on dumps server labstore1007 for the weekend because it has been flapping [[phab:T281045|T281045]] === 2021-05-13 === * 21:25 bstorm: converted the maps and scratch volumes on cloudstore1008 (standby) to drbd [[phab:T224747|T224747]] * 15:45 bstorm: re-running wikireplicas-dns after refactor of config to make sure it doesn't change anything === 2021-05-12 === * 14:23 arturo: [codfw1dev] cleanup old unused agents (bgp, ovs) * 11:37 arturo: [codfw1dev] replacing cloudnet2003-dev with cloudnet2004-dev ([[phab:T281381|T281381]]) === 2021-05-11 === * 18:00 andrewbogott: adding 'trove' service project in advance of deploying trove in eqiad1 * 10:22 arturo: rebooted cloudgw1002 (active) thus causing a failover to cloudgw1001 === 2021-05-09 === * 10:53 arturo: icinga-downtime cloudmetrics1002 for 3 months ([[phab:T275605|T275605]]) === 2021-05-07 === * 13:51 andrewbogott: add inherited 'admin' right to novaadmin user throughout eqiad1. I was trying to narrow down the rights here but lack of admin breaks some workflows, e.g. [[phab:T281894|T281894]] and [[phab:T282235|T282235]] === 2021-05-06 === * 15:31 arturo: about to migrating CloudVPS network to the cloudgw architecture [[phab:T270704|T270704]] * 11:14 dcaro: restarting cinder-volume on the eqiad control nodes to refresh the ceph libraries ([[phab:T282109|T282109]]) === 2021-05-05 === * 16:07 dcaro: disallowing insecure global ids on the eqiad ceph cluster ([[phab:T280641|T280641]]) * 15:15 wm-bot: Safe reboot of 'cloudvirt1046.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:11 wm-bot: Safe rebooting 'cloudvirt1046.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:11 wm-bot: Safe reboot of 'cloudvirt1045.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:07 wm-bot: Safe rebooting 'cloudvirt1045.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:07 wm-bot: Safe reboot of 'cloudvirt1044.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:03 wm-bot: Safe rebooting 'cloudvirt1044.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:03 wm-bot: Safe reboot of 'cloudvirt1043.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 14:59 wm-bot: Safe rebooting 'cloudvirt1043.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 14:59 wm-bot: Safe reboot of 'cloudvirt1042.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 14:40 wm-bot: Safe rebooting 'cloudvirt1042.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 14:39 wm-bot: Safe reboot of 'cloudvirt1041.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 14:14 wm-bot: Safe rebooting 'cloudvirt1041.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 14:14 wm-bot: Safe reboot of 'cloudvirt1039.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 14:10 wm-bot: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 12:35 wm-bot: Safe rebooting 'cloudvirt1039.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 11:56 wm-bot: Safe rebooting 'cloudvirt1038.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 11:56 wm-bot: Safe reboot of 'cloudvirt1037.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 11:31 wm-bot: Safe rebooting 'cloudvirt1037.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 11:31 wm-bot: Safe reboot of 'cloudvirt1036.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 11:08 wm-bot: Safe rebooting 'cloudvirt1036.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 11:08 wm-bot: Safe reboot of 'cloudvirt1035.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 10:39 wm-bot: Safe rebooting 'cloudvirt1035.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 10:39 wm-bot: Safe reboot of 'cloudvirt1034.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 10:13 wm-bot: Safe rebooting 'cloudvirt1034.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 10:13 wm-bot: Safe reboot of 'cloudvirt1033.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 09:47 wm-bot: Safe rebooting 'cloudvirt1033.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 09:47 wm-bot: Safe reboot of 'cloudvirt1032.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 09:21 wm-bot: Safe rebooting 'cloudvirt1032.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 09:21 wm-bot: Safe reboot of 'cloudvirt1031.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:45 wm-bot: Safe rebooting 'cloudvirt1031.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:45 wm-bot: Safe reboot of 'cloudvirt1030.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:19 wm-bot: Safe rebooting 'cloudvirt1030.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:19 wm-bot: Safe reboot of 'cloudvirt1029.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:02 wm-bot: Safe rebooting 'cloudvirt1029.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus === 2021-05-04 === * 16:05 wm-bot: Safe reboot of 'cloudvirt1028.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:45 wm-bot: Safe rebooting 'cloudvirt1028.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:44 wm-bot: Safe reboot of 'cloudvirt1027.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:22 wm-bot: Safe rebooting 'cloudvirt1027.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:19 wm-bot: Safe reboot of 'cloudvirt1026.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:15 wm-bot: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 13:19 dcaro: rebooting cloudmetrics1002, got stuck again ([[phab:T275605|T275605]]) * 10:04 wm-bot: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 09:10 wm-bot: Safe rebooting 'cloudvirt1026.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 09:10 wm-bot: Safe reboot of 'cloudvirt1025.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:34 wm-bot: Safe rebooting 'cloudvirt1025.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:20 wm-bot: Safe reboot of 'cloudvirt1024.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 08:03 wm-bot: Safe rebooting 'cloudvirt1024.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus === 2021-05-03 === * 23:53 bstorm: running `maintain-dbusers harvest-replicas` on labstore1004 [[phab:T281287|T281287]] * 23:51 bstorm: running `maintain-dbusers harvest-replicas` on labstore1004 * 16:34 wm-bot: Safe reboot of 'cloudvirt1023.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 16:29 wm-bot: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:41 wm-bot: Safe rebooting 'cloudvirt1023.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:41 wm-bot: Safe reboot of 'cloudvirt1022.eqiad.wmnet' finished successfully. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 15:13 wm-bot: Safe rebooting 'cloudvirt1022.eqiad.wmnet'. ([[phab:T280641|T280641]]) - cookbook ran by dcaro@vulcanus * 10:31 wm-bot: Safe rebooting 'cloudvirt1021.eqiad.wmnet'. ([[phab:T280641|T280641]] - cookbook ran by dcaro@vulcanus) * 10:23 wm-bot: (from a cookbook) * 09:12 dcaro: draining and rebooting coludvirt1021 ([[phab:T280641|T280641]]) * 08:26 dcaro: draining and rebooting coludvirt1018 ([[phab:T280641|T280641]]) === 2021-04-30 === * 11:16 dcaro: draining and rebooting coludvirt1017, last one today ([[phab:T280641|T280641]]) * 10:37 dcaro: draining coludvirt1016 for reboot ([[phab:T280641|T280641]]) * 09:48 dcaro: draining coludvirt1013 for reboot ([[phab:T280641|T280641]]) === 2021-04-29 === * 15:11 dcaro: hard rebooting cloudmetrics1002, got hung again ([[phab:T275605|T275605]]) * 07:53 dcaro: Upgrading ceph libraries on cloudcontrol1005 to octopus ([[phab:T274566|T274566]]) * 07:51 dcaro: Upgrading ceph libraries on cloudcontrol1003 to octopus ([[phab:T274566|T274566]]) * 07:50 dcaro: Upgrading ceph libraries on cloudcontrol1004 to octopus ([[phab:T274566|T274566]]) === 2021-04-28 === * 21:11 andrewbogott: cleaning up more references to deleted hypervisors with delete from services where topic='compute' and version != 53; * 20:48 andrewbogott: cleaning up references to deleted hypervisors with mysql:root@localhost [nova_eqiad1]> delete from compute_nodes where hypervisor_version != '5002000'; * 19:40 andrewbogott: putting cloudvirt1040 into the maintenance aggregate pending more info about [[phab:T281399|T281399]] * 18:11 andrewbogott: adding cloudvirt1040, 1041 and 1042 to the 'ceph' host aggregate -- [[phab:T275081|T275081]] * 11:06 dcaro: All ceph server side upgraded to Octopus! \o/ ([[phab:T280641|T280641]]) * 10:57 dcaro: Got a PG getting stuck on 'remapping' after the OSD came up, had to unset the norebalance and then set it again to get it unstuck ([[phab:T280641|T280641]]) * 10:34 dcaro: Slow/blocked opns from cloudcephmon03, "osd_failure(failed timeout osd.32..." (cloudcephosd1005), unset the cluster noout/norebalance and went away in a few secs, setting it again and continuing... ([[phab:T280641|T280641]]) * 09:03 dcaro: Waiting for slow heartbeats from osd.58(cloudcephosd1002) to recover... ([[phab:T280641|T280641]]) * 08:59 dcaro: During the upgrade, started getting warning 'slow osd heartbacks in the back', meaning that pings between osds are really slow (up to 190s) all from osd.58, currently on cloudcephosd1002 ([[phab:T280641|T280641]]) * 08:58 dcaro: During the upgrade, started getting warning 'slow osd heartbacks in the back', meaning that pings between osds are really slow (up to 190s) all from osd.58 ([[phab:T280641|T280641]]) * 08:58 dcaro: During the upgrade, started getting warning 'slow osd heartbacks in the back', meaning that pings between osds are really slow (up to 190s) ([[phab:T280641|T280641]]) * 08:21 dcaro: Upgrading all the ceph osds on eqiad ([[phab:T280641|T280641]]) * 08:21 dcaro: The clock skew seems intermittent, there's another task to follw it [[phab:T275860|T275860]] ([[phab:T280641|T280641]]) * 08:18 dcaro: All equiad ceph mons and mgrs upgraded ([[phab:T280641|T280641]]) * 08:18 dcaro: During the upgrade, ceph detected a clock skew on cloudcephmon1002, cloudcephmon1001, they are back ([[phab:T280641|T280641]]) * 08:15 dcaro: During the upgrade, ceph detected a clock skew on cloudcephmon1002, it went away, I'm guessing systemd-timesyncd fixed it ([[phab:T280641|T280641]]) * 08:14 dcaro: During the upgrade, ceph detected a clock skew on cloudcephmon1002, looking ([[phab:T280641|T280641]]) * 07:58 dcaro: Upgrading ceph services on eqiad, starting with mons/managers ([[phab:T280641|T280641]]) === 2021-04-27 === * 14:10 dcaro: codfw.openstack upgraded ceph libraries to 15.2.11 ([[phab:T280641|T280641]]) * 13:07 dcaro: codfw.openstack cloudvirt2002-dev done, taking cloudvirt2003-dev out to upgrade ceph libraries ([[phab:T280641|T280641]]) * 13:00 dcaro: codfw.openstack cloudvirt2001-dev back online, taking cloudvirt2002-dev out to upgrade ceph libraries ([[phab:T280641|T280641]]) * 10:51 dcaro: ceph.eqiad: cinder pool got it's pg_num increased to 1024, re-shuffle started ([[phab:T273783|T273783]]) * 10:48 dcaro: ceph.eqiad: Tweaked the target_size_ratio of all the pools, enabling autoscaler (it will increase cinder pool only) ([[phab:T273783|T273783]]) * 09:14 dcaro: manually force stopping the server puppetmaster-01 to unblock migration (in codfw1) * 09:14 dcaro: manually force stopping the server puppetmaster-01 to unblock migration * 08:59 dcaro: manually force stopping the server exploding-head on codfw, to try cold migration * 08:47 dcaro: restarting nova-compute on cloudvirt2001-dev after upgrading ceph libraries to 15.2.11 === 2021-04-26 === * 20:56 andrewbogott: deleting spurious 'codfw1dev' and 'codw1dev-4' regions in the dallas deployment; regions without endpoints break a bunch of things * 09:45 dcaro: draining cloudvirt2001-dev with the new cookbooks ([[phab:T280641|T280641]]) === 2021-04-23 === * 13:49 dcaro: testing the drain_cloudvirt cookbook on codfw1 openstack cluster, draining cloudvirt2001 ([[phab:T280641|T280641]]) * 11:12 dcaro: testing the drain_cloudvirt cookbook on codfw1 openstack cluster ([[phab:T280641|T280641]]) * 09:32 dcaro: finished upgrade of ceph cluster on codfw1 using exclusively cookbooks ([[phab:T280641|T280641]]) * 09:17 dcaro: testing the upgrade_osds cookbook on codfw1 ceph cluster ([[phab:T280641|T280641]]) * 08:17 dcaro: testing the upgrade_mons cookbook on codfw1 ceph cluster ([[phab:T280641|T280641]]) === 2021-04-21 === * 17:59 dcaro: all monitors upgraded on codfw1 with one cookbook `cookbook --verbose -c ~/.config/spicerack/cookbook.yaml wmcs.ceph.upgrade_mons --monitor-node-fqdn cloudcephmon2002-dev.codfw.wmnet` ([[phab:T280641|T280641]]) * 17:47 dcaro: upgrading monitors and mrg nodes on codfw ceph cluster ([[phab:T280641|T280641]]) * 13:26 dcaro: testing ceph upgrade cookbook on cloudcephmon2002-dev ([[phab:T280641|T280641]]) === 2021-04-20 === * 20:21 andrewbogott: reboot cloudservices1003 * 20:13 andrewbogott: reboot cloudservices1004 === 2021-04-19 === * 08:40 dcaro: enabling puppet on labstore1004 after mysql restart ([[phab:T279657|T279657]]) * 08:09 dcaro: downtiming labstore1004 and stopping puppet for mysql restart ([[phab:T279657|T279657]]) === 2021-04-14 === * 10:48 dcaro: Upgrade of codfw ceph to octopus 15.2.20 done, will run some performance tests now ([[phab:T274566|T274566]]) * 10:41 dcaro: Upgrade of codfw ceph to octopus 15.2.20, mgrs upgraded, osds next ([[phab:T274566|T274566]]) * 10:37 dcaro: Upgrade of codfw ceph to octopus 15.2.20, mons upgraded, mgrs next ([[phab:T274566|T274566]]) * 10:15 dcaro: starting the upgrade of codfw ceph to octopus 15.2.20 ([[phab:T274566|T274566]]) * 10:07 dcaro: Merged the ceph 15 (Octopus) repo deployment to codfw, only the repo, not the packages ([[phab:T274566|T274566]]) === 2021-04-13 === * 16:42 dcaro: Ceph balancer got the cluster to eval 0.014916, that is 88-77% usage for compute pool, and 28-19% usage for the cinder one \o/ ([[phab:T274573|T274573]]) * 15:08 dcaro: Activating continuous upmap balancer, keeping a close eye ([[phab:T274573|T274573]]) * 15:03 dcaro: Executing a second pass, there's still movements to improve the eval of 0.030075 ([[phab:T274573|T274573]]) * 15:02 dcaro: First pass finished, improved eval to 0.030075 ([[phab:T274573|T274573]]) * 14:49 dcaro: Running the first_pass balancing plan on ceph eqiad, current eval 0.030622 ([[phab:T274573|T274573]]) * 14:43 dcaro: enabling ceph upmap pg balancer on equiad ([[phab:T274573|T274573]]) * 14:36 andrewbogott: upgrading codfw1dev to version Victoria, [[phab:T261137|T261137]] * 13:11 andrewbogott: upgrading eqiad1 designate to version Victoria, [[phab:T261137|T261137]] * 10:44 dcaro: enabled ceph upmap balancer on codfw ([[phab:T274573|T274573]],[[phab:T274573|T274573]]) === 2021-04-07 === * 21:33 andrewbogott: upgrading codfw1dev designate to Victoria === 2021-04-04 === * 17:36 andrewbogott: upgrading eqiad1 designate to Ussuri === 2021-04-02 === * 14:12 andrewbogott: upgrading codfw1dev to OpenStack version Ussuri === 2021-04-01 === * 12:15 dcaro: Restoring the 4.9 kernel on cloudcephosd2003-dev and upgrading ([[phab:T274565|T274565]]) * 10:29 dcaro: Done restoring the 4.9 kernel on cloudcephosd2001-dev and upgrading, requires logging into console to boot from the older kernel before removing the newer one ([[phab:T274565|T274565]]) * 10:10 dcaro: Restoring the 4.9 kernel on cloudcephosd2001-dev and upgrading ([[phab:T274565|T274565]]) === 2021-03-31 === * 08:47 dcaro: upgrading cinder on codfw cloudcontrol2* nodes ([[phab:T278845|T278845]]) === 2021-03-30 === * 09:53 arturo: rebooting cloudnet1003 to cleanup conntrack table, it wouldn't cleanup by hand ... === 2021-03-28 === * 15:42 andrewbogott: updated debian-10.0-buster base image === 2021-03-27 === * 09:54 arturo: cleanup conntrack table in qrouter nents in cloudnet1003 (backup) === 2021-03-25 === * 19:03 andrewbogott: deleting all unused (per wmcs-imageusage) Jessie base images from Glance * 17:15 andrewbogott: refreshing puppet compiler facts for tools project * 10:31 dcaro: kernel upgrade on osds on codfw done, running performance tests ([[phab:T274565|T274565]]) * 10:24 dcaro: upgrading kernel on cloudcephosd2003-dev and reboot ([[phab:T274565|T274565]]) * 10:18 dcaro: upgrading kernel on cloudcephosd2002-dev and reboot ([[phab:T274565|T274565]]) * 10:08 dcaro: upgrading kernel on cloudcephmon2003-dev and reboot ([[phab:T274565|T274565]]) === 2021-03-24 === * 09:19 dcaro: restarted wmcs-backup on cloudvirt1024 as it failed due to an image being removed while running ([[phab:T276892|T276892]]) === 2021-03-23 === * 11:33 arturo: root@cloudcontrol1005:~# wmcs-novastats-dnsleaks --delete === 2021-03-22 === * 10:10 arturo: cleanup conntrack table in standby node: aborrero@cloudnet1003:~ $ sudo ip netns exec qrouter-d93771ba-2711-4f88-804a-{{Gerrit|8df6fd03978a}} conntrack -F === 2021-03-19 === * 17:18 bstorm: running `ALTER TABLE account MODIFY COLUMN type ENUM('user','tool','paws');` against the labsdbaccounts database on m5 [[phab:T276284|T276284]] * 14:29 andrewbogott: switching admin-monitoring project to use an upstream debian image; I want to see how this affects performance * 00:30 bstorm: downtimed labstore1004 to check some things in debug mode === 2021-03-17 === * 17:28 bstorm: restarted the backup-glance-images job to clear errors in systemd [[phab:T271782|T271782]] * 17:16 andrewbogott: set default cinder quota for projects to 80Gb with "update quota_classes set hard_limit=80 where resource='gigabytes';" on database 'cinder' * 16:58 andrewbogott: disabling all flavors with >20Gb root storage with "update flavors set disabled=1 where root_gb>20;" in nova_eqiad1_api === 2021-03-10 === * 16:51 arturo: rebooting cloudvirt1030 for [[phab:T275753|T275753]] * 13:14 dcaro: starting manually the canary VM for cloudvirt1029 (nova start 349830f6-3b39-4a8c-ada4-{{Gerrit|a7439f65cffe}}) ([[phab:T275753|T275753]]) * 12:51 arturo: draining cloudvirt1030 for [[phab:T275753|T275753]] * 12:47 arturo: rebooting cloudvirt1029 for [[phab:T275753|T275753]] * 11:56 arturo: [codfw1dev] restart rabbitmq-server in all 3 cloudcontrol servers for [[phab:T276964|T276964]] * 11:53 arturo: [codfw1dev] restart nova-conductor in all 3 cloudcontrol servers for [[phab:T276964|T276964]] * 11:31 arturo: draining cloudvirt1029 for [[phab:T275753|T275753]] * 11:29 arturo: rebooting cloudvirt1013 for [[phab:T275753|T275753]] * 11:05 arturo: draining cloudvirt1013 for [[phab:T275753|T275753]] * 11:00 arturo: rebooting cloudvirt1028 for [[phab:T275753|T275753]] * 10:33 arturo: draining cloudvirt1028 for [[phab:T275753|T275753]] * 10:29 arturo: rebooting cloudvirt1023 for [[phab:T275753|T275753]] * 09:37 arturo: draining cloudvirt1023 for [[phab:T275753|T275753]] * 09:07 arturo: [codfw1dev] reimaging cloudvirt2003-dev ([[phab:T276964|T276964]]) === 2021-03-09 === * 16:27 arturo: rebooting cloudvirt1027 ([[phab:T275753|T275753]]) * 13:39 arturo: draining cloudvrit1027 for [[phab:T275753|T275753]] * 13:35 arturo: icinga-downtime cloudvirt1038 for 30 days for [[phab:T276922|T276922]] * 13:21 arturo: add cloudvirt1039 to the ceph host aggregate (no longer a spare, we have cloudvirt1038 with HW failures) * 12:52 arturo: cloudvirt1038 hard powerdown / powerup for [[phab:T276922|T276922]] * 12:33 arturo: rebooting cloudvirt1038 ([[phab:T275753|T275753]]) * 10:58 arturo: draining cloudvirt1038 ([[phab:T275753|T275753]]) * 10:54 arturo: rebooting cloudvirt1037 ([[phab:T275753|T275753]]) * 09:59 arturo: draining cloudvirt1037 ([[phab:T275753|T275753]]) * 09:12 dcaro: restarted the wmcs-backup service on cloudvirt1024 to retry the backups (failed because a VM was removed in-between, [[phab:T276892|T276892]]) === 2021-03-05 === * 21:40 andrewbogott: replacing 'observer' role with 'reader' role in eqiad1 [[phab:T276018|T276018]] * 21:21 andrewbogott: replacing 'observer' role with 'reader' role in eqiad1 * 16:23 arturo: rebooting cloudvirt1036 for [[phab:T275753|T275753]] * 12:30 arturo: draining cloudvirt1036 for [[phab:T275753|T275753]] * 12:25 arturo: rebooting cloudvirt1035 for [[phab:T275753|T275753]] * 10:49 arturo: rebooting cloudvirt1035 for [[phab:T275753|T275753]] * 10:47 arturo: rebooting cloudvirt1034 for [[phab:T275753|T275753]] * 10:26 arturo: draining cloudvirt1034 for [[phab:T275753|T275753]] * 10:25 arturo: rebooting cloudvirt1033 for [[phab:T275753|T275753]] * 09:18 arturo: draining cloudvirt1033 for [[phab:T275753|T275753]] === 2021-03-04 === * 18:36 andrewbogott: rebooting cloudmetrics1002; the console is hanging * 16:59 arturo: rebooting cloudvirt1032 for [[phab:T275753|T275753]] * 16:34 arturo: draining cloudvirt1032 for [[phab:T275753|T275753]] * 16:33 arturo: rebooting cloudvirt1031 for [[phab:T275753|T275753]] * 16:11 arturo: draining cloudvirt1031 for [[phab:T275753|T275753]] * 16:09 arturo: rebooting cloudvirt1026 for [[phab:T275753|T275753]] * 15:57 arturo: draining cloudvirt1026 for [[phab:T275753|T275753]] * 15:55 arturo: rebooting cloudvirt1025 for [[phab:T275753|T275753]] * 15:41 arturo: draining cloudvirt1025 for [[phab:T275753|T275753]] * 15:12 arturo: rebooting cloudvirt1024 for [[phab:T275753|T275753]] * 11:29 arturo: draining cloudvirt1024 for [[phab:T275753|T275753]] * 11:24 dcaro: rebooted cloudvirt1022, re-adding to ceph and removing from maintenance host aggregate for [[phab:T275753|T275753]] * 11:01 dcaro: rebooting cloudvirt1022 for [[phab:T275753|T275753]] * 09:12 dcaro: draining cloudvirt1022 for [[phab:T275753|T275753]] === 2021-03-03 === * 17:16 andrewbogott: restarting rabbitmq-server on cloudcontrol1003,1004,1005; trying to explain amqp errors in scheduler logs * 16:03 dcaro: draining cloudvirt1022 for [[phab:T275753|T275753]] * 16:03 dcaro: draining cloudvirt1022 for [[phab:T275753|T275753]] * 16:00 arturo: move cloudvirt1013 into the 'toobusy' host aggregate, it has 221% cpu subscription and 82% MEM subscription * 15:34 arturo: rebooting cloudvirt1021 for [[phab:T275753|T275753]] * 14:31 arturo: draining cloudvirt1021 for [[phab:T275753|T275753]] * 13:59 arturo: rebooting cloudvirt1018 for [[phab:T275753|T275753]] * 13:28 arturo: draining cloudvirt1018 for [[phab:T275753|T275753]] * 12:49 arturo: rebooting cloudvirt1017 for [[phab:T275753|T275753]] * 12:22 arturo: draining cloudvirt1017 for [[phab:T275753|T275753]] * 12:20 arturo: rebooting cloudvirt1016 for [[phab:T275753|T275753]] * 12:01 arturo: draining cloudvirt1016 for [[phab:T275753|T275753]] * 11:59 arturo: cloudvirt1014 now in the ceph host aggregate * 11:58 arturo: rebooting cloudvirt1014 for [[phab:T275753|T275753]] * 11:50 arturo: moved cloudvirt1023 away from the maintenance host aggregate, leave it in the ceph aggregate (was in the 2) * 11:47 arturo: moved cloudvirt1014 to the 'maintenance' host aggregate, drain it for [[phab:T275753|T275753]] * 10:01 arturo: icinga-downtime cloudnet1003 for 14 days bc potential alerting storm due to firmware issues ([[phab:T271058|T271058]]) * 10:01 arturo: rebooting again cloudnet1003 (no network failover) ([[phab:T271058|T271058]]) * 09:59 arturo: update firmware-bnx2x from 20190114-2 to 20200918-1~bpo10+1 on cloudnet1003 ([[phab:T271058|T271058]]) * 09:30 arturo: installing linux kernel 5.10.13-1~bpo10+1 in cloudnet1003 and rebooting it (network failover) ([[phab:T271058|T271058]]) === 2021-03-02 === * 17:16 andrewbogott: rebooting cloudvirt1039 to see if I can trigger [[phab:T276208|T276208]] * 16:10 arturo: [codfw1dev] restart nova-compute on cloudvirt2002-dev * 11:59 arturo: moved cloudvirt1012 to 'maintenance' host aggregate. Drain it with `wmcs-drain-hypervisor` to reboot it for [[phab:T275753|T275753]] * 11:59 arturo: cloudvirt1023 is affected by [[phab:T276208|T276208]] and cannot be rebooted. Put it back into the ceph hos aggregate * 10:43 arturo: moved cloudvirt1013 cloudvirt1032 cloudvirt1037 back into the 'ceph' host aggregate * 10:13 arturo: moved cloudvirt1023 to 'maintenance' host aggregate. Drain it with `wmcs-drain-hypervisor` to reboot it for [[phab:T275753|T275753]] === 2021-03-01 === * 20:12 andrewbogott: removing novaadmin from all projects save 'admin' for [[phab:T274385|T274385]] * 19:51 andrewbogott: removing novaobserver from all projects save 'observer' for [[phab:T274385|T274385]] * 19:50 andrewbogott: adding inherited domain-wide roles to novaadmin and novaobserver as per [[phab:T274385|T274385]] === 2021-02-28 === * 04:54 andrewbogott: restarted redis-server on tools-redis-1003 and tools-redis-1004 in an attempt to reduce replag, no real change detected === 2021-02-27 === * 00:33 andrewbogott: sudo cumin --timeout 500 "A:all and not O<nowiki>{</nowiki>project:clouddb-services<nowiki>}</nowiki>" 'lsb_release -c {{!}} grep -i buster && uname -r {{!}} grep -v 4.19.0-14-amd64 && reboot' * 00:28 andrewbogott: sudo cumin --timeout 500 "A:all and not O<nowiki>{</nowiki>project:clouddb-services<nowiki>}</nowiki>" 'lsb_release -c {{!}} grep -i buster && uname -r {{!}} grep -v 4.19.0-14-amd64 && echo reboot' * 00:09 andrewbogott: sudo cumin "A:all and not O<nowiki>{</nowiki>project:clouddb-services<nowiki>}</nowiki>" 'lsb_release -c {{!}} grep -i stretch && uname -r {{!}} grep -v 4.19.0-0.bpo.14-amd64 && reboot' === 2021-02-26 === * 14:58 dcaro: [eqiad] rebooting cloudcephosd1015 (last osd \o/) for kernel upgrade ([[phab:T275753|T275753]]) * 14:51 dcaro: [eqiad] rebooting cloudcephosd1014 for kernel upgrade ([[phab:T275753|T275753]]) * 14:44 dcaro: [eqiad] rebooting cloudcephosd1013 for kernel upgrade ([[phab:T275753|T275753]]) * 14:38 dcaro: [eqiad] rebooting cloudcephosd1012 for kernel upgrade ([[phab:T275753|T275753]]) * 14:31 dcaro: [eqiad] rebooting cloudcephosd1011 for kernel upgrade ([[phab:T275753|T275753]]) * 14:25 dcaro: [eqiad] rebooting cloudcephosd1010 for kernel upgrade ([[phab:T275753|T275753]]) * 14:17 dcaro: [eqiad] rebooting cloudcephosd1009 for kernel upgrade ([[phab:T275753|T275753]]) * 13:54 dcaro: [eqiad] downtimed alert1001 Ceph OSDs down alert until 18:00 GMT+1 as that is not under the host being rebooted ([[phab:T275753|T275753]]) * 13:51 dcaro: [eqiad] rebooting cloudcephosd1008 for kernel upgrade ([[phab:T275753|T275753]]) * 13:45 dcaro: [eqiad] rebooting cloudcephosd1007 for kernel upgrade ([[phab:T275753|T275753]]) * 13:38 dcaro: [eqiad] rebooting cloudcephosd1006 for kernel upgrade ([[phab:T275753|T275753]]) * 12:07 dcaro: [eqiad] rebooting cloudcephosd1005 for kernel upgrade ([[phab:T275753|T275753]]) * 12:00 arturo: rebooting cloudcontrol1003 for kernel upgrade ([[phab:T275753|T275753]]) * 11:42 arturo: rebooting cloudcontrol1004 for kernel upgrade ([[phab:T275753|T275753]]) * 11:41 dcaro: [eqiad] rebooting cloudcephosd1004 for kernel upgrade ([[phab:T275753|T275753]]) * 11:32 dcaro: [eqiad] rebooting cloudcephosd1003 for kernel upgrade ([[phab:T275753|T275753]]) * 11:30 arturo: rebooting cloudcontrol1005 for kernel upgrade ([[phab:T2|T2]] * 11:26 dcaro: [eqiad] rebooting cloudcephosd1002 for kernel upgrade ([[phab:T275753|T275753]]) * 11:16 dcaro: [eqiad] rebooting cloudcephosd1001 for kernel upgrade ([[phab:T275753|T275753]]) * 11:11 dcaro: [eqiad] rebooting cloudcephmon1003 for kernel upgrade ([[phab:T275753|T275753]]) * 11:05 dcaro: [eqiad] rebooting cloudcephmon1002 for kernel upgrade ([[phab:T275753|T275753]]) * 10:59 dcaro: [eqiad] rebooting cloudcephmon1001 for kernel upgrade ([[phab:T275753|T275753]]) * 10:45 arturo: rebooting cloudvirt1039 into a new kernel ([[phab:T275753|T275753]]) --- spare * 10:43 dcaro: [codfw1dev] rebooting cloudcephmon2003-dev for kernel upgrade ([[phab:T275753|T275753]]) * 10:38 dcaro: [codfw1dev] rebooting cloudcephmon2002-dev for kernel upgrade ([[phab:T275753|T275753]]) * 10:29 dcaro: [codfw1dev] rebooting cloudcephmon2001-dev for kernel upgrade ([[phab:T275753|T275753]]) * 10:24 arturo: [codfw1dev] purge old kernel packages on cloudvirt2003-dev to force boot into a new kernel ([[phab:T275753|T275753]]) * 10:11 arturo: [codfw1dev] manually creating /boot/grub/ on cloudvirt2003-dev to allow update-grub2 to run (so it can reboot into a new kernel) ([[phab:T275753|T275753]]) * 10:11 dcaro: [codfw1dev] rebooting cloudcephosd2003-dev for kernel upgrade ([[phab:T275753|T275753]]) * 10:05 dcaro: [codfw1dev] rebooting cloudcephosd2002-dev for kernel upgrade ([[phab:T275753|T275753]]) * 10:01 arturo: [codfw1dev] rebooting cloudvirt200X-dev for kernel upgrade ([[phab:T275753|T275753]]) * 09:59 arturo: [codfw1dev] rebooting cloudweb2001-dev for kernel upgrade ([[phab:T275753|T275753]]) * 09:53 arturo: [codfw1dev] rebooting cloudservices2003-dev for kernel upgrade ([[phab:T275753|T275753]]) * 09:51 arturo: [codfw1dev] rebooting cloudservices2002-dev for kernel upgrade ([[phab:T275753|T275753]]) * 09:45 arturo: [codfw1dev] rebooting cloudcontrol2004-dev for kernel upgrade ([[phab:T275753|T275753]]) * 09:44 arturo: [codfw1dev] rebooting cloudbackup[2001-2002].codfw.wmnet for kernel upgrade ([[phab:T275753|T275753]]) * 09:43 dcaro: [codfw1dev] rebooting cloudcephosd2001-dev for kernel upgrade ([[phab:T275753|T275753]]) * 09:41 arturo: [codfw1dev] rebooting cloudcontrol2003-dev for kernel upgrade ([[phab:T275753|T275753]]) * 09:33 arturo: [codfw1dev] rebooting cloudcontrol2001-dev for kernel upgrade ([[phab:T275753|T275753]]) === 2021-02-25 === * 14:56 arturo: deployed wmcs-netns-events daemon to all cloudnet servers ([[phab:T275483|T275483]]) === 2021-02-24 === * 11:07 arturo: force-reboot cloudmetrics1002, add icinga downtime for 2 hours. Investigating some server issue * 00:17 bstorm: set --property hw_scsi_model=virtio-scsi and --property hw_disk_bus=scsi on the main stretch image in glance on eqiad1 [[phab:T275430|T275430]] === 2021-02-23 === * 22:43 bstorm: set --property hw_scsi_model=virtio-scsi and --property hw_disk_bus=scsi on the main buster image in glance on eqiad1 [[phab:T275430|T275430]] * 20:36 andrewbogott: adding r/o access to the eqiad1-glance-images ceph pool for the client.eqiad1-compute for [[phab:T275430|T275430]] * 10:49 arturo: rebooting clounet1004 into new kernel from buster-bpo ([[phab:T271058|T271058]]) * 10:49 arturo: installing linux-image-amd64 from buster-bpo 5.10.13-1~bpo10+1 in cloudnet1004 ([[phab:T271058|T271058]]) === 2021-02-22 === * 17:15 bstorm: restarting nova-compute on cloudvirt1016 and cloudvirt1036 in case it helps [[phab:T275411|T275411]] * 15:02 dcaro: Re-uploaded the debian buster 10.0 image from rbd to glance, that worked, re-spawning all the broken instances ([[phab:T275378|T275378]]) * 11:12 dcaro: Refreshing all the canary instances ([[phab:T275354|T275354]]) === 2021-02-18 === * 14:50 arturo: rebooting cloudnet1004 for [[phab:T271058|T271058]] * 10:25 dcaro: Rebooting cloudmetrics1001 to apply new kernel ([[phab:T275116|T275116]]) * 10:16 dcaro: Rebooting cloudmetrics1002 to apply new kernel ([[phab:T275116|T275116]]) * 10:14 dcaro: Upgrading grafana on cloudmetrics1002 ([[phab:T275116|T275116]]) * 10:12 dcaro: Upgrading grafana on cloudmetrics1001 ([[phab:T275116|T275116]]) === 2021-02-17 === * 15:58 arturo: deploying https://gerrit.wikimedia.org/r/c/operations/puppet/+/664845 to cloudnet servers ([[phab:T268335|T268335]]) === 2021-02-15 === * 16:25 arturo: [codfw1dev] rebooting all cloudgw200x-dev / cloudnet200x-dev servers ([[phab:T272963|T272963]]) * 15:45 arturo: [codfw1dev] drop subnet definition for cloud-instances-transport1-b-codfw ([[phab:T272963|T272963]]) * 15:45 arturo: [codfw1dev] connect virtual router cloudinstances2b-gw to vlan cloud-gw-transport-codfw (185.15.57.10) ([[phab:T272963|T272963]]) === 2021-02-11 === * 12:01 arturo: [codfw1dev] drop instance `tools-codfw1dev-bastion-1` in `tools-codfw1dev` (was buster, cannot use it yet) * 11:59 arturo: [codfw1dev] create instance `tools-codfw1dev-bastion-2` (stretch) in `tools-codfw1dev` to test stuff related to [[phab:T272397|T272397]] * 11:45 arturo: [codfw1dev] create instance `tools-codfw1dev-bastion-1` in `tools-codfw1dev` to test stuff related to [[phab:T272397|T272397]] * 11:42 arturo: [codfw1dev] drop `tools` project, create `tools-codfw1dev` * 11:38 arturo: [codfw1dev] drop `coudinfra` project (we are using `cloudinfra-codfw1dev` there) * 05:37 bstorm: downtimed cloudnet1004 for another week [[phab:T271058|T271058]] === 2021-02-09 === * 15:23 arturo: icinga-downtime for 2h everything *labs *cloud for openstack upgrades * 11:14 dcaro: Merged the osd scheduler change for all osds, applying on all cloudcephosd* ([[phab:T273791|T273791]]) === 2021-02-08 === * 18:50 bstorm: enabled puppet on cloudvirt1023 for now [[phab:T274144|T274144]] * 18:44 bstorm: restarted the backup_vms.service on cloudvirt1027 [[phab:T274144|T274144]] * 17:51 bstorm: deleted project pki [[phab:T273175|T273175]] === 2021-02-05 === * 10:59 arturo: icinga-downtime labstore1004 tools share space check for 1 week ([[phab:T272247|T272247]]) * 10:21 dcaro: This was affecting maps and several others, maps and project-proxy have been fixed ([[phab:T273956|T273956]]) * 09:19 dcaro: Some certs around the infra are expired ([[phab:T273956|T273956]]) === 2021-02-04 === * 10:12 dcaro: Increasing the memory limit of osds in eqiad from 8589934592(8G) to 12884901888(12G) ([[phab:T273851|T273851]]) === 2021-02-03 === * 09:59 dcaro: Doing a full vm backup on cloudvirt1024 with the new script ([[phab:T260692|T260692]]) * 01:50 bstorm: icinga-downtime cloudnet1004 for a week [[phab:T271058|T271058]] === 2021-02-02 === * 17:14 dcaro: Changed osd memory limit from 4G to 8G ([[phab:T273649|T273649]]) * 11:00 arturo: icinga-downtime cloudvirt-wdqs1001 for 1 week ([[phab:T273579|T273579]]) * 03:12 andrewbogott: running /usr/local/sbin/wmcs-purge-backups and /usr/local/sbin/wmcs-backup-instances on cloudvirt1024 to see why the backup job paged === 2021-01-29 === * 15:36 andrewbogott: disabling puppet and some services on eqiad1 cloudcontrol nodes; replacing nova-placement-api with placement-api === 2021-01-28 === * 19:44 andrewbogott: shutting down cloudcontrol2001-dev because it's in a partially upgraded state; will revive when it's time for Train === 2021-01-27 === * 00:50 bstorm: icinga-downtime cloudnet1004 for a week [[phab:T271058|T271058]] === 2021-01-22 === * 16:44 andrewbogott: upgrading designate on cloudvirt1003/1004 to OpenStack 'train' * 11:29 dcaro: Doing some tests removed cloudcontrol1003 puppet cert, regenerating... === 2021-01-21 === * 11:35 arturo: merging core router firewall changes https://gerrit.wikimedia.org/r/c/operations/homer/public/+/657439 ([[phab:T209082|T209082]]) * 11:30 arturo: merging core router firewall changes https://gerrit.wikimedia.org/r/c/operations/homer/public/+/657358 ([[phab:T272486|T272486]], [[phab:T209082|T209082]]) === 2021-01-20 === * 10:49 arturo: merging core router firewall change https://gerrit.wikimedia.org/r/c/operations/homer/public/+/657302 ([[phab:T209082|T209082]]) * 10:05 dcaro: Everything looks ok, created a new vm with a volume in ceph without issues, and on warnings/errors on ceph status, closing ([[phab:T272303|T272303]]) * 09:55 dcaro: Eqiad ceph cluster uprgaded, doing sanity checks ([[phab:T272303|T272303]]) * 09:46 dcaro: 75% of the eqiad cluster upgraded... continuing ([[phab:T272303|T272303]]) * 09:37 dcaro: 25% of the eqiad cluster upgraded... continuing ([[phab:T272303|T272303]]) * 09:24 dcaro: Mgr daemons upgraded and running, upgrading osd daemons on servers cloudcephosd1*, this make take a bit longer ([[phab:T272303|T272303]]) * 09:22 dcaro: Mon daemons upgraded and running, upgrading mgr daemons on servers cloudcephmon1* ([[phab:T272303|T272303]]) * 09:16 dcaro: Starting eqiad ceph upgrade, upgrading the mon servers cloudcephmon1* ([[phab:T272303|T272303]]) * 09:01 dcaro: Will start the ceph upgrade in 15 min, no downtime nor performance impact is expected ([[phab:T272303|T272303]]) === 2021-01-19 === * 10:17 arturo: icinga-downtime cloudnet1004 for 1 week ([[phab:T271058|T271058]]) === 2021-01-18 === * 16:00 dcaro: Codfw1 ceph cluster uprgaded, will wait until tomorrow to see if there's any instability, but everything looks fine ([[phab:T272303|T272303]]) * 15:38 dcaro: Upgraded mgr sevices on codfw ceph cluster, starting with osd ones ([[phab:T272303|T272303]]) * 15:35 dcaro: Upgraded mon sevices on codfw ceph cluster, starting with mgr ones ([[phab:T272303|T272303]]) * 15:21 dcaro: Starting upgrade of ceph mon nodes on codfw ([[phab:T272303|T272303]]) * 15:06 dcaro: re-enabling puppet on cloudcephosd2* hosts * 13:53 dcaro: disabling puppet on cloudcephosd2* to resume perf tests * 10:50 dcaro: re-enabling puppet on cephcloudosd2* (codfw) * 10:07 dcaro: disabling puppet on cephcloudosd2* (codfw) to do some performance tests * 09:00 dcaro: Enabling custom application 'cinder' on pool codfw1dev-cinder to get rid of health warnings === 2021-01-17 === * 16:53 arturo: icinga downtime labstore1004 /srv/tools space check for 3 days ([[phab:T272247|T272247]]) === 2021-01-15 === * 13:41 arturo: icinga downtime labstore1004 maintain-dbuser alert until 2021-01-19 ([[phab:T272125|T272125]]) * 09:47 arturo: labstore1004 maintain-dbusers affected by [[phab:T272127|T272127]] and [[phab:T272125|T272125]] * 09:22 arturo: restart maintain-dbusers.service in labstore1004 * 08:19 dcaro: Merging the patch to disable write caches on ceph osds ([[phab:T271527|T271527]]) === 2021-01-13 === * 17:03 arturo: remove cloudvirt1013 cloudvirt1032 cloudvirt1037 to the 'toobusy' host aggregate to prevent further CPU oversubscribing * 12:40 arturo: try increasing systemd watchdog timeout for conntrackd in cloudnet1004 ([[phab:T268335|T268335]]) * 11:45 dcaro: https://gerrit.wikimedia.org/r/c/operations/puppet/+/654419 merged and deployed (and tested) ([[phab:T268877|T268877]]) * 11:40 dcaro: merging https://gerrit.wikimedia.org/r/c/operations/puppet/+/654419 that might affect the encapi service (puppet on cloud environment), no downtime expected though ([[phab:T268877|T268877]]) * 10:56 arturo: trying to cleanup dpkg package mess in cloudnet2002-dev * 10:02 arturo: prevent floating IP allocation from neutron transport subnet: root@cloudcontrol1005:~# neutron subnet-update --allocation-pool start=185.15.56.244,end=185.15.56.244 cloud-instances-transport1-b-eqiad1 ([[phab:T271867|T271867]]) === 2021-01-12 === * 10:33 arturo: reboot cloudnet1004 * 10:32 arturo: update firmware-bnx2x from 20190114-2 to 20200918-1~bpo10+1 on cloudnet1004 ([[phab:T271058|T271058]]) === 2021-01-11 === * 10:22 arturo: doubling size of conntrack table in cloudnet servers https://gerrit.wikimedia.org/r/c/operations/puppet/+/655407 ([[phab:T271058|T271058]]) * 10:07 arturo: manually cleanup conntrack table in cloudnet1004 ([[phab:T271058|T271058]]) * 09:19 dcaro: cleaned up ~1800 snapshots, 109 remaining only, one for each host x image combination (plus some ephemeral ones while doing backups), closing the task ([[phab:T270478|T270478]]) * 08:39 dcaro: cleaning up dangling snapshots now that we have the new suffixed ones ([[phab:T270478|T270478]]) === 2021-01-10 === * 16:02 andrewbogott: restarting rabbitmq-server on all eqiad1 cloudcontrols * 15:54 andrewbogott: restating neutron-metadata-agent on cloudnet1004 due to many syslog complaints === 2021-01-08 === * 11:25 arturo: rebooting both cloudnet2002-dev/cloudnet2003-dev to make sure interfaces are set up correctl ([[phab:T271517|T271517]]) * 11:22 arturo: connecting cloudnet2002-dev cloudnet2003-dev back to vlan 2120 ([[phab:T271517|T271517]]) * 11:06 arturo: root@cloudcontrol2001-dev:~# openstack router set --external-gateway wan-transport-codfw --fixed-ip subnet=cloud-instances-transport1-b-codfw,ip-address=208.80.153.190 cloudinstances2b-gw ([[phab:T271517|T271517]]) * 11:02 arturo: root@cloudcontrol2001-dev:~# openstack router set --enable-snat cloudinstances2b-gw --external-gateway wan-transport-codfw ([[phab:T271517|T271517]]) * 11:01 arturo: enabling neutron hacks in codfw1dev (cloudnet2002-dev, cloudnet2003-dev) ([[phab:T271517|T271517]]) * 10:55 arturo: aborrero@labtestvirt2003:~ $ sudo ifdown eno2.2107 ([[phab:T271517|T271517]]) * 10:55 arturo: aborrero@labtestvirt2003:~ $ sudo ifdown eno2.2120 ([[phab:T271517|T271517]]) * 10:53 arturo: root@cloudcontrol2001-dev:~# openstack subnet create --network wan-transport-codfw --gateway 208.80.153.185 --ip-version 4 --network wan-transport-codfw --no-dhcp --subnet-range 208.80.153.184/29 cloud-instances-transport1-b-codfw ([[phab:T271517|T271517]]) * 10:40 dcaro: Finished tests, brining osd online (od.48) for eqiad ceph cluster ([[phab:T271417|T271417]]) * 09:59 dcaro: Started performance tests on sdc (od.48) for eqiad ceph cluster ([[phab:T271417|T271417]]) * 09:41 dcaro: Taking osd.48 from eqiad ceph cluster out to do performance tests ([[phab:T271417|T271417]]) === 2021-01-07 === * 15:19 dcaro: Finished speed tests on cloudcephosd2001-dev, reprovisioning the osd.0 sdc ([[phab:T271417|T271417]]) * 14:39 dcaro: Starting speed tests on cloudcephosd2001-dev sdc ([[phab:T271417|T271417]]) * 12:54 dcaro: Taking osd.0 down on codfw ceph cluster to try the disk performance testing process ([[phab:T271417|T271417]]) * 11:35 arturo: merging dmz_cidr change ([[phab:T209082|T209082]], [[phab:T267779|T267779]]) === 2021-01-05 === * 10:40 dcaro: removing dumps-[1..*] backups from cloudvirt1024 as they are not needed ([[phab:T271094|T271094]]) === 2021-01-03 === * 07:06 dcaro: Got a network hiccup on cloudnet1004, keeping track here [[phab:T271058|T271058]] === 2020-12-28 === * 12:32 arturo: stop doing backups for the dumps project https://gerrit.wikimedia.org/r/c/operations/puppet/+/652182 ([[phab:T260692|T260692]]) * 12:32 arturo: stop doing backups for the dumps project https://gerrit.wikimedia.org/r/c/operations/puppet/+/652182 ([[phab:T260682|T260682]]) * 12:23 arturo: icinga downtime cloudvirt1026 disk space check until january 5 ([[phab:T260692|T260692]]) * 06:15 andrewbogott: restarting designate-central on cloudservices1003/1004. I'm pretty sure they're distressed because of DB lag but it's worth a try === 2020-12-23 === * 15:38 andrewbogott: restarting rabbitmq on cloudcontrol1004; suspected leaks * 15:33 andrewbogott: restarting each cloudcontrol galera node in turn to see if that quiets down the syncing warnings * 12:08 arturo: move memory out of the swap in cloudcontrol1004 by disabling/enabling it (1Gb swap was being used) === 2020-12-22 === * 15:30 dcaro: cleaning up 6778 dangling snapshots for glance images in eqiad ([[phab:T270478|T270478]]) * 13:51 dcaro: merged patch to move wikidumpparse backups to cloudvirt1025 to free space on cloudvirt1026 === 2020-12-19 === * 16:18 dcaro: gzipped a bunch of logs on cloudvirt1004 due to / being out of space * 00:14 bstorm: truncated /var/log/debug.1 on cloudcontrol1003 which appears to be the exact same content as the user.log files anyway * 00:10 bstorm: truncated /var/log/daemon.log.1 and the haproxy log * 00:02 bstorm: truncated /var/log/messages.1 on cloudcontrol1003 === 2020-12-18 === * 23:53 bstorm: truncated haproxy.log.1 on cloudcontrol1003 * 20:46 andrewbogott: setting pg and pgp number to 4096 for eqiad1-compute as joachim thinks 8192 might be too much [[phab:T270305|T270305]] * 17:09 dcaro: finished cleaning up the dangling snapshots from cloudvirt1026 ([[phab:T270478|T270478]]) * 17:08 dcaro: removing dangling rbd snapshots (for backups on cloudvirt1026) ([[phab:T270478|T270478]]) * 17:06 dcaro: finished cleaning up the dangling snapshots from cloudvirt1025 ([[phab:T270478|T270478]]) * 17:05 dcaro: removing dangling rbd snapshots (for backups on cloudvirt1025) ([[phab:T270478|T270478]]) * 17:00 dcaro: finished cleaning up the dangling snapshots from cloudvirt1021 ([[phab:T270478|T270478]]) * 16:58 dcaro: removing dangling rbd snapshots (for backups on cloudvirt1021) ([[phab:T270478|T270478]]) * 16:56 dcaro: finished cleaning up the dangling snapshots from cloudvirt1022 ([[phab:T270478|T270478]]) * 16:55 dcaro: removing dangling rbd snapshots (for backups on cloudvirt1022) ([[phab:T270478|T270478]]) * 16:54 dcaro: finished cleaning up the dangling snapshots from cloudvirt1023 ([[phab:T270478|T270478]]) * 16:51 dcaro: removing dangling rbd snapshots (for backups on cloudvirt1023) ([[phab:T270478|T270478]]) * 16:47 dcaro: finished cleaning up the dangling snapshots from cloudvirt1024, freed ~12% of the capacity ([[phab:T270478|T270478]]) * 16:21 dcaro: removing dangling rbd snapshots (for backups on cloudvirt1024) ([[phab:T270478|T270478]]) * 16:13 andrewbogott: setting autoscale to 'off' for both ceph pools (eqiad1-compute and eqiad1-glance-images) because we like how things are set and the autoscaler does not * 10:33 dcaro: purging rbd snapshots for image fc6fb78b-4515-4dcc-8254-{{Gerrit|591b9fe01762}} ([[phab:T270478|T270478]]) === 2020-12-17 === * 22:17 andrewbogott: correction to above, set the pg and pgp to 1024 for eqiad1-glance-images * 22:16 andrewbogott: setting pgp number to 8192 for eqiad1-compute (a 4x increase) and 2048 for eqiad1-glance-images (also a 4x increase) [[phab:T270305|T270305]] (same as pg) * 22:14 andrewbogott: setting pg number to 8192 for eqiad1-compute (a 4x increase) and 2048 for eqiad1-glance-images (also a 4x increase) [[phab:T270305|T270305]] * 22:10 andrewbogott: setting autoscale to 'warn' for both ceph pools (eqiad1-compute and eqiad1-glance-images) === 2020-12-16 === * 09:31 dcaro: removing invalid backups from cloudvirt1024 (196 in total) ([[phab:T269419|T269419]]) === 2020-12-14 === * 17:42 dcaro: The removal freed ~12GB (still 100% usage :S) ([[phab:T269419|T269419]]) * 17:36 dcaro: removing invalid backups that have a valid copy ([[phab:T269419|T269419]]) * 15:43 dcaro: Merging the tagging for vm backups ([[phab:T267195|T267195]]) * 09:45 arturo: icinga downtime cloudvirt1024 for 6 days ([[phab:T269419|T269419]]) === 2020-12-13 === * 09:11 _dcaro: running backup purge script on cloudvirt1024 ([[phab:T269419|T269419]]) === 2020-12-10 === * 23:36 bstorm: cleaned up the logs for haproxy on cloudcontrol1003 by deleting all the gzipped ones and truncating the .1 file * 11:56 dcaro: Freed some space on cloudvirt1024 by running the purge script ([[phab:T269419|T269419]]) * 09:17 dcaro: removing leaked dns record discordwiki.eqiad.wmflabs (clinic duty) === 2020-12-08 === * 18:01 dcaro: Host cloudvirt1030 up and running ([[phab:T216195|T216195]]) * 15:59 dcaro: Re-imaging host cloudvirt1030 ([[phab:T216195|T216195]]) * 14:18 dcaro: Host online cloudvirt1029 ([[phab:T216195|T216195]]) * 14:13 dcaro: Host re-imaged, doing tests cloudvirt1029 ([[phab:T216195|T216195]]) * 12:14 dcaro: Re-imaging cloudvirt1029 ([[phab:T216195|T216195]]) === 2020-12-07 === * 18:33 andrewbogott: putting cloudvirt1023 back into service [[phab:T269467|T269467]] * 15:55 andrewbogott: reimaging cloudvirt1028 for [[phab:T216195|T216195]] * 14:49 dcaro: Re-imaging cloudvirt1027 ([[phab:T216195|T216195]]) === 2020-12-05 === * 00:35 andrewbogott: moving cloudvirt1023 back into maintenance because [[phab:T269467|T269467]] continues to puzzle === 2020-12-04 === * 22:33 andrewbogott: moving cloudvirt1023 back into the ceph aggregate; it doesn't need upgrades after all [[phab:T269467|T269467]] * 22:24 andrewbogott: moving cloudvirt1023 out of the ceph aggregate and into maintenance for [[phab:T269467|T269467]] * 21:06 andrewbogott: putting cloudvirt1025 and 1026 back into service because I'm pretty sure they're fixed. [[phab:T269313|T269313]] * 12:12 arturo: manually running `wmcs-purge-backups` again on cloudvirt1024 ([[phab:T269419|T269419]]) * 11:25 arturo: icinga downtime cloudvirt1024 for 6 days, to avoid paging noises ([[phab:T269419|T269419]]) * 11:25 arturo: last log line referencing cloudvirt1024 is a mistake ([[phab:T269313|T269313]]) * 11:24 arturo: icinga downtime cloudvirt1024 for 6 days, to avoid paging noises ([[phab:T269313|T269313]]) * 10:28 arturo: manually running `wmcs-purge-backups` on cloudvirt1024 ([[phab:T269419|T269419]]) * 10:23 arturo: setting expiration to 2020-12-03 to the oldest backy snapshot of every VM in cloudvirt1024 ([[phab:T269419|T269419]]) * 09:54 arturo: icinga downtime cloudvirt1025 for 6 days ([[phab:T269313|T269313]]) === 2020-12-03 === * 23:21 andrewbogott: removing all osds on cloudcephosd1004 for rebuild, [[phab:T268746|T268746]] * 21:45 andrewbogott: removing all osds on cloudcephosd1005 for rebuild, [[phab:T268746|T268746]] * 19:51 andrewbogott: removing all osds on cloudcephosd1006 for rebuild, [[phab:T268746|T268746]] * 17:01 arturo: icinga downtime cloudvirt1025 for 48h to debug network issue [[phab:T269313|T269313]] * 16:56 arturo: rebooting cloudvirt1025 to debug network issue [[phab:T269313|T269313]] * 16:38 dcaro: Rimaging cloudvirt1026 ([[phab:T216195|T216195]]) * 13:24 andrewbogott: removing all osds on cloudcephosd1008 for rebuild, [[phab:T268746|T268746]] * 02:55 andrewbogott: removing all osds on cloudcephosd1009 for rebuild, [[phab:T268746|T268746]] === 2020-12-02 === * 20:04 andrewbogott: removing all osds on cloudcephosd1010 for rebuild, [[phab:T268746|T268746]] * 17:25 arturo: [15:51] failovering neutron virtual router in eqiad1 ([[phab:T268335|T268335]]) * 15:36 arturo: conntrackd is now up and running in cloudnet1003/1004 nodes ([[phab:T268335|T268335]]) * 15:33 arturo: [codfw1dev] conntrackd is now up and running in cloudnet200x-dev nodes ([[phab:T268335|T268335]]) * 15:08 andrewbogott: removing all osds on cloudcephosd1012 for rebuild, [[phab:T268746|T268746]] * 12:41 arturo: disable puppet in all cloudnet servers to merge conntrackd change [[phab:T268335|T268335]] * 11:12 dcaro: Reset the properties for the flavor g2.cores8.ram16.disk1120 to correct quotes ([[phab:T269172|T269172]]) * 09:57 arturo: moved cloudvirts 1030, 1029, 1028, 1027, 1026, 1025 away from the 'standard' host aggregate to 'maintenance' ([[phab:T269172|T269172]]) === 2020-12-01 === * 20:06 andrewbogott: removing all osds on cloudcephosd1014 for rebuild, [[phab:T268746|T268746]] * 12:04 arturo: restarting neutron l3 agents to pick up config change * 11:48 arturo: merging change to dmz_dir, detail list of private address https://gerrit.wikimedia.org/r/c/operations/puppet/+/641977 === 2020-11-30 === * 18:12 andrewbogott: removing all osds from cloudcephosd1015 in order to investigate [[phab:T268746|T268746]] === 2020-11-29 === * 17:18 andrewbogott: cleaning up some logfiles in tools-sgecron-01 — drive is full === 2020-11-26 === * 22:58 andrewbogott: deleting /var/log/haproxy logs older than 7 days in cloudcontrol100x. We need log rotation here it seems. * 15:53 dcaro: Created private flavor g2.cores8.ram16.disk1120 for wikidumpparse ([[phab:T268190|T268190]]) === 2020-11-25 === * 19:35 bstorm: repairing ceph pg `instructing pg 6.91 on osd.117 to repair` * 09:31 _dcaro: The OSD seems to be up and running actually, though there's that misleading log, will leave it see if the cluster comes fully healthy ([[phab:T268722|T268722]]) * 08:54 _dcaro: Unsetting noup/nodown to allow re-shuffling of the pgs that osd.44 had, will try to rebuild it ([[phab:T268722|T268722]]) * 08:45 _dcaro: Tried resetting the class for osd.44 to ssd, no luck, the cluster is in noout/norebalance to avoid data shuffling (opened [[phab:T268722|T268722]]) * 08:45 _dcaro: Tried resetting the class for osd.44 to ssd, no luck, the cluster is in noout/norebalance to avoid data shuffling (opened root@cloudcephosd1005:/var/lib/ceph/osd/ceph-44# ceph osd crush set-device-class ssd osd.44) * 08:19 _dcaro: Restarting serivce osd.44 resulted on osd.44 being unable to start due to some config inconsistency (can not reset class to hdd) * 08:16 _dcaro: After enabling auto pg scaling on ceph eqiad cluster, osd.44 (cloudcephosd1005) got stuck, trying to restart the osd service * 08:16 _dcaro: After enabling auto pg scaling on ceph eqiad cluster, osd.44 (cloudcephosd1005) got stuck, trying to restart === 2020-11-22 === * 17:40 andrewbogott: apt-get upgrade on cloudservices1003/1004 * 17:32 andrewbogott: upgrading Designate on cloudservices1003/1004 to Stein === 2020-11-20 === * 12:44 arturo: [codfw1dev] install conntrackd in cloudnet2003-dev/cloudnet2002-dev to research l3 agent HA reliability * 09:26 arturo: incinga downtime labstore1006 RAID checks for 10 days ([[phab:T268281|T268281]]) === 2020-11-17 === * 19:21 andrewbogott: draining cloudvirt1012 to experiment with libvirt/cpu things === 2020-11-15 === * 11:21 arturo: icinga downtime cloudbackup2002 for 48h ([[phab:T267865|T267865]]) === 2020-11-10 === * 16:38 arturo: icinga downtime toolschecker for 2h becasue toolsdb maintenance ([[phab:T266587|T266587]]) * 11:24 arturo: [codfw1dev] enable puppet in puppetmaster01.cloudinfra-codfw1dev (disabled for unspecified reasons) === 2020-11-09 === * 12:42 arturo: restarted neutron l3 agent in cloudnet1003 bc it still had the old default route ([[phab:T265288|T265288]]) * 12:41 arturo: `root@cloudcontrol1005:~# neutron subnet-delete dcbb0f98-5e9d-4a93-8dfc-4e3ec3c44dcc` ([[phab:T265288|T265288]]) * 12:41 arturo: `root@cloudcontrol1005:~# neutron router-gateway-set --fixed-ip subnet_id=7c6bcc12-212f-44c2-9954-{{Gerrit|5c55002ee371}},ip_address=185.15.56.244 cloudinstances2b-gw wan-transport-eqiad` ([[phab:T265288|T265288]]) * 12:19 arturo: subnet 185.1.5.56.240/29 has id 7c6bcc12-212f-44c2-9954-{{Gerrit|5c55002ee371}} in neutron ([[phab:T265288|T265288]]) * 12:19 arturo: `root@cloudcontrol1005:~# neutron subnet-create --gateway 185.15.56.241 --name cloud-instances-transport1-b-eqiad1 --ip-version 4 --disable-dhcp wan-transport-eqiad 185.15.56.240/29` ([[phab:T265288|T265288]]) * 12:15 arturo: icinga-downtime toolschecker for 2h ([[phab:T265288|T265288]]) === 2020-11-02 === * 13:36 arturo: (typo: dcaro) * 13:35 arturo: added dcar as projectadmin & user ([[phab:T266068|T266068]]) === 2020-10-29 === * 16:57 bstorm: silenced deployment-prep project alerts for 60 days since the downtime expired * 08:12 arturo: force-powercycling cloudcephosd1006 === 2020-10-25 === * 16:20 andrewbogott: adding cloudvirt1038 to the 'ceph' aggregate and removing from the 'spare' aggregate. We need this space while waiting on network upgrades for empty cloudvirts ([[phab:T216195|T216195]]) === 2020-10-23 === * 11:30 arturo: [codfw1dev] openstack --os-project-id cloudinfra-codfw1dev recordset create --type PTR --record nat.cloudgw.codfw1dev.wikimediacloud.org. --description "created by hand" 0-29.57.15.185.in-addr.arpa. 1.0-29.57.15.185.in-addr.arpa. ([[phab:T261724|T261724]]) * 10:09 arturo: [codf1dev] doing DNS changes for the cloudgw PoC, including designate and https://gerrit.wikimedia.org/r/c/operations/dns/+/635965 ([[phab:T261724|T261724]]) === 2020-10-22 === * 10:46 arturo: [codfw1dev] rebooting cloudinfra-internal-puppetmaster-01.cloudinfra-codfw1dev.codfw1dev.wikimedia.cloud to try fixing some DNS weirdness * 09:43 arturo: enabling puppet in cloucontrol1003 (message said "please re-enable after 2020-10-22 06:00UTC") === 2020-10-21 === * 14:36 andrewbogott: running apt-get update && apt-get install -y facter on all cloud-vps instances * 10:31 arturo: [codfw1dev] reimaging labtestvirt2003 (cloudgw) to test puppet code ([[phab:T261724|T261724]]) * 08:56 arturo: [codfw1dev] reimaging labtestvirt2003 (cloudgw) to test puppet code ([[phab:T261724|T261724]]) === 2020-10-20 === * 15:47 arturo: changing DNS recursor ACLs (https://gerrit.wikimedia.org/r/c/operations/puppet/+/635314) this can be reverted any time if it causes problems ([[phab:T261724|T261724]]) * 14:49 arturo: [codfw1dev] reimaging labtestvirt2003 (cloudgw) to test puppet code ([[phab:T261724|T261724]]) === 2020-10-19 === * 01:41 andrewbogott: deleting all Precise base images * 01:36 andrewbogott: deleting all unused Jessie base images === 2020-10-18 === * 23:26 andrewbogott: deleting all Trusty base images * 21:50 andrewbogott: migrating all currently used ceph images to rbd === 2020-10-16 === * 09:29 arturo: [codfw1dev] still some DNS weirdness, investigating * 09:25 arturo: [codfw1dev] hard-rebooting bastion-codfw1dev-02, seems in bad shape, doesn't even wake up in the virsh console * 09:18 arturo: [codfw1dev] live-hacked cloudservices2002-dev /etc/powerdns/recursor.conf file to include cloud-codfw1dev-floating CIDR (185.15.57.0/29) while https://gerrit.wikimedia.org/r/c/operations/puppet/+/634050 is in review, so VMs with a floating IP can query the DNS recursor ([[phab:T261724|T261724]]) * 09:01 arturo: [codfw1dev] basic network connectivity seems stable after cleaning up everything related to address scopes ([[phab:T261724|T261724]]) === 2020-10-15 === * 15:17 arturo: [codfw1dev] try cleaning up anything related to address scopes in the neutron database ([[phab:T261724|T261724]]) * 13:56 arturo: [codfw1dev] drop neutron l3 agent hacks in cloudnet2002/2003-dev ([[phab:T261724|T261724]]) === 2020-10-13 === * 17:54 andrewbogott: rebuilding cloudvirt1021 for backy support * 15:22 andrewbogott: draining cloudvirt1021 so I can rebuild it with backy support * 14:19 andrewbogott: rebuilding cloudvirt1022 with backy support * 14:03 andrewbogott: draining cloudvirt1022 so I can rebuild it with backy support * 11:19 arturo: [codfw1dev] rebooting labtestvirt2003 === 2020-10-09 === * 10:15 arturo: [codfwd1ev] root@cloudcontrol2001-dev:~# openstack router set --disable-snat cloudinstances2b-gw --external-gateway wan-transport-codfw ([[phab:T261724|T261724]]) * 09:22 arturo: [codfwd1dev] rebooting cloudnet boxes for bridge and vlan changes ([[phab:T261724|T261724]]) * 09:12 arturo: [codfw1dev] root@cloudcontrol2001-dev:~# openstack subnet delete 31214392-9ca5-4256-bff5-{{Gerrit|1e19a35661de}} (cloud-instances-transport1-b-codfw - 208.80.153.184/29) ([[phab:T261724|T261724]]) * 09:10 arturo: [codfw1dev] root@cloudcontrol2001-dev:~# openstack router set --external-gateway wan-transport-codfw --fixed-ip subnet=cloud-gw-transport-codfw,ip-address=185.15.57.10 cloudinstances2b-gw ([[phab:T261724|T261724]]) * 08:49 arturo: [codfw1dev] root@cloudcontrol2001-dev:~# openstack subnet create --network wan-transport-codfw --gateway 185.15.57.9 --no-dhcp --subnet-range 185.15.57.8/30 cloud-gw-transport-codfw ([[phab:T261724|T261724]]) * 08:47 arturo: [codfw1dev] root@cloudcontrol2001-dev:~# openstack subnet delete a5ab5362-4ffb-4059-9ff7-{{Gerrit|391e22dcf3bc}} ([[phab:T261724|T261724]]) === 2020-10-08 === * 16:17 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# openstack subnet create --network wan-transport-codfw --gateway 185.15.57.8 --no-dhcp --subnet-range 185.15.57.8/31 cloud-gw-transport-codfw` (with a hack -- see task) ([[phab:T263622|T263622]]) * 16:03 arturo: [codfw1dev] briefly live-hacked python3-neutron source code in all 3 cloudcontrol2xxx-dev servers to workaround /31 network definition issue ([[phab:T263622|T263622]]) * 10:28 arturo: [codfw1dev] reimaging labtestvirt2003 (cloudgw) [[phab:T261724|T261724]] === 2020-10-06 === * 21:30 andrewbogott: moved cloudvirt1013 out of the 'ceph' aggregate and into the 'maintenance' aggregate for [[phab:T243414|T243414]] * 21:29 andrewbogott: draining cloudvirt1013 for upgrade to 10G networking * 14:45 arturo: icinga downtime every cloud* lab* host for 60 minutes for keystone maintenance === 2020-10-05 === * 17:40 bd808: `service uwsgi-labspuppetbackend restart` on cloud-puppetmaster-03 ([[phab:T264649|T264649]]) === 2020-10-02 === * 11:05 arturo: [codfw1dev] restarting rabbitmq-server in all 3 control nodes, the l3 agent was misbehaving * 09:16 arturo: [codfw1dev] trying the labtestvirt2003 (cloudgw) reimage again ([[phab:T261724|T261724]]) === 2020-10-01 === * 16:06 arturo: rebooting cloudvirt1024 to validate changes to /etc/network/interfaces file * 15:36 arturo: [codfw1dev] reimaging labtestvirt2003 === 2020-09-30 === * 16:47 andrewbogott: rebooting cloudvir1032, 1033, 1034 for [[phab:T262979|T262979]] * 13:28 arturo: enable puppet, reboot and pool back cloudvirt1031 * 13:27 arturo: extend icinga downtimes for another 120 mins * 13:15 arturo: `aborrero@cloudcontrol1003:~$ sudo nova-manage placement sync_aggregates` after reading a hint in nova-api.log * 13:02 arturo: rebooting cloudvirt1016 and moving it to the ceph host aggregate * 12:55 arturo: rebooting cloudvirt1014 and moving it to the ceph host aggregate * 12:51 arturo: rebooting cloudvirt1013 and moving it to the ceph host aggregate * 12:39 arturo: root@cloudcontrol1005:~# openstack aggregate add host maintenance cloudvirt1031 * 12:36 arturo: rebooted cloudnet1003 (active) a couple of minutes ago * 12:36 arturo: move cloudvirt1012 and cloudvirt1039 to the ceph aggregate * 11:49 arturo: rebooting cloudvirt1039 * 11:46 arturo: rebooting cloudvirt1012 * 11:40 arturo: rebooting cloudnet1004 (standby) to pick up https://gerrit.wikimedia.org/r/c/operations/puppet/+/631167 ([[phab:T262979|T262979]]) * 11:38 arturo: [codfw1dev] rebooting cloudnet2002-dev to pick up https://gerrit.wikimedia.org/r/c/operations/puppet/+/631167 * 11:36 arturo: [codfw1dev] rebooting cloudnet2003-dev to pick up https://gerrit.wikimedia.org/r/c/operations/puppet/+/631167 * 11:33 arturo: disabling puppet and downtiming every virt/net server in the fleet in preparation for merging https://gerrit.wikimedia.org/r/c/operations/puppet/+/631167 ([[phab:T262979|T262979]]) * 09:32 arturo: rebooting cloudvirt1012 to investigate linuxbridge agent issues === 2020-09-29 === * 15:40 arturo: downgrade linux kernel from linux-image-4.19.0-11-amd64 to linux-image-4.19.0-10-amd64 on cloudvirt1012 * 14:47 arturo: rebooting cloudvirt1012, chasing config weirdness in the linuxbridge agent * 14:05 andrewbogott: reimaging 1014 over and over in an attempt to get partman right * 13:51 arturo: rebooting cloudvirt1012 === 2020-09-28 === * 14:55 arturo: [jbond42] upgraded facter to v3 across the VM fleet * 13:54 andrewbogott: moving cloudvirt1035 from aggregate 'spare' to 'ceph'. We're going to need all the capacity we can get while converting older cloudvirts to ceph === 2020-09-24 === * 15:47 arturo: stopping/restarting rabbitmq-server in all cloudcontrol servers * 15:45 arturo: restarting rabbitmq-server in cloudcontrol103 * 15:15 arturo: restarting floating_ip_ptr_records_updater.service in all 3 cloudcontrol servers to reset state after a DNS failure === 2020-09-18 === * 10:16 arturo: cloudvirt1039 libvirtd service issues were fixed with a reboot * 09:56 arturo: rebooting cloudvirt1039 (spare) to try to fix some weird libvirtd failure * 09:50 arturo: enabling puppet in cloudvirts and effectively merging patches from [[phab:T262979|T262979]] * 08:59 arturo: disable puppet in all buster cloudvirts (cloudvirt[1024,1031-1039].eqiad.wmnet) to merge a patch for [[phab:T263205|T263205]] and [[phab:T262979|T262979]] * 08:50 arturo: installing iptables from buster-bpo in cloudvirt1036 ([[phab:T263205|T263205]] and [[phab:T262979|T262979]]) === 2020-09-15 === * 20:32 andrewbogott: rebooting cloudvirt1038 to see if it resolves [[phab:T262979|T262979]] * 13:58 andrewbogott: draining cloudvirt1002 with wmcs-ceph-migrate === 2020-09-14 === * 14:21 andrewbogott: draining cloudvirt1001, migrating all VMs with wmcs-ceph-migrate * 10:41 arturo: [codfw1dev] trying to get the bonding working for labtestvirt2003 ([[phab:T261724|T261724]]) * 09:47 arturo: installed qemu security update in eqiad1 cloudvirts ([[phab:T262386|T262386]]) * 09:43 arturo: [codfw1dev] installed qemu security update in codfw1dev cloudvirts ([[phab:T262386|T262386]]) === 2020-09-09 === * 18:13 andrewbogott: restarting ceph-mon@cloudcephmon1003 in hopes that the slow ops reported are phantoms * 18:01 andrewbogott: restarting ceph-mgr@cloudcephmon1003 in hopes that the slow ops reported are phantoms (https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/thread/EOWNO3MDYRUZKAK6RMQBQ5WBPQNLHOPV/) * 17:40 andrewbogott: giving ceph pg autoscale another chance: ceph osd pool set eqiad1-compute pg_autoscale_mode on * 00:05 bd808: Running wmcs-novastats-dnsleaks ([[phab:T262359|T262359]]) === 2020-09-08 === * 21:48 bd808: Renamed FQDN prefixes to wikimedia.cloud scheme in cloudinfra-db01's labspuppet db ([[phab:T260614|T260614]]) * 14:29 andrewbogott: restarting nova-compute on all cloudvirts (everyone is upset from the reset switch failure) * 14:18 arturo: restarting nova-fullstack service in cloudcontrol1003 * 14:17 andrewbogott: stopping apache2 on labweb1001 to make sure the Horizon outage is total === 2020-09-03 === * 09:31 arturo: icinga downtime cloud* servers for 30 mins ([[phab:T261866|T261866]]) === 2020-09-02 === * 08:46 arturo: [codfw1dev] reimaging spare server labtestvirt2003 as debian buster ([[phab:T261724|T261724]]) === 2020-09-01 === * 18:18 andrewbogott: adding drives on cloudcephosd100[3-5] to ceph osd pool * 13:40 andrewbogott: adding drives on cloudcephosd101[0-2] to ceph osd pool * 13:35 andrewbogott: adding drives on cloudcephosd100[1-3] to ceph osd pool * 11:27 arturo: [codfw1dev] rebooting again cloudnet2002-dev after some network tests, to reset initial state ([[phab:T261724|T261724]]) * 11:09 arturo: [codfw1dev] rebooting cloudnet2002-dev after some network tests, to reset initial state ([[phab:T261724|T261724]]) * 10:49 arturo: disable puppet in cloudnet servers to merge https://gerrit.wikimedia.org/r/c/operations/puppet/+/623569/ === 2020-08-31 === * 23:26 bd808: Removed stale lockfile at cloud-puppetmaster-03.cloudinfra.eqiad.wmflabs:/var/lib/puppet/volatile/GeoIP/.geoipupdate.lock * 11:20 arturo: [codfw1dev] livehacking https://gerrit.wikimedia.org/r/c/operations/puppet/+/615161 in the puppetmasters for tests before merging === 2020-08-28 === * 20:12 bd808: Running `wmcs-novastats-dnsleaks --delete` from cloudcontrol1003 === 2020-08-26 === * 17:12 bstorm: Running 'ionice -c 3 nice -19 find /srv/tools -type f -size +100M -printf "%k KB %p\n" > tools_large_files_20200826.txt' on labstore1004 [[phab:T261336|T261336]] === 2020-08-21 === * 21:34 andrewbogott: restarting nova-compute on cloudvirt1033; it seems stuck === 2020-08-19 === * 14:21 andrewbogott: rebooting cloudweb2001-dev, labweb1001, labweb1002 to address mediawiki-induced memleak === 2020-08-06 === * 21:02 andrewbogott: removing cloudvirt1004/1006 from nova's list of hypervisors; rebuilding them to use as backup test hosts * 20:06 bstorm: manually stopped the RAID check on cloudcontrol1003 [[phab:T259760|T259760]] === 2020-08-04 === * 18:54 bstorm: restarting mariadb on cloudcontrol1004 to setup parallel replication === 2020-08-03 === * 17:02 bstorm: increased db connection limit to 800 across galera cluster because we were clearly hovering at limit === 2020-07-31 === * 19:28 bd808: wmcs-novastats-dnsleaks --delete (lots of leaked fullstack-monitoring records to clean up) === 2020-07-27 === * 22:17 andrewbogott: ceph osd pool set compute pg_num 2048 * 22:14 andrewbogott: ceph osd pool set compute pg_autoscale_mode off === 2020-07-24 === * 19:15 andrewbogott: ceph mgr module enable pg_autoscaler * 19:15 andrewbogott: ceph osd pool set compute pg_autoscale_mode on === 2020-07-22 === * 08:55 jbond42: [codfw1dev] upgrading hiera to version5 * 08:48 arturo: [codfw1dev] add jbond as user in the bastion-codfw1dev and cloudinfra-codfw1dev projects * 08:45 arturo: [codfw1dev] enabled account creation in labtestwiki briefly for jbond42 to create an account === 2020-07-16 === * 10:48 arturo: merging change to neutron dmz_cidr https://gerrit.wikimedia.org/r/c/operations/puppet/+/613123 ([[phab:T257534|T257534]]) === 2020-07-15 === * 23:15 bd808: Removed Merlijn van Deen from toollabs-trusted Gerrit group ([[phab:T255697|T255697]]) * 11:48 arturo: [codfw1dev] created DNS records (A and PTR) for bastion.bastioninfra-codfw1dev.codfw1dev.wmcloud.org <-> 185.15.57.2 * 11:41 arturo: [codfw1dev] add myself as projectadmin to the `bastioninfra-codfw1dev` project * 11:39 arturo: [codfw1dev] created DNS zone `bastioninfra-codfw1dev.codfw1dev.wmcloud.org.` in the cloudinfra-codfw1dev project and then transfer ownership to the bastioninfra-codfw1dev project === 2020-07-14 === * 15:19 arturo: briefly set root@cloudnet1003:~ # sysctl net.ipv4.conf.all.accept_local=1 (in neutron qrouter netns) ([[phab:T257534|T257534]]) * 10:43 arturo: icinga downtime cloudnet* hosts for 30 mins to introduce new check https://gerrit.wikimedia.org/r/c/operations/puppet/+/612390 ([[phab:T257552|T257552]]) * 04:01 andrewbogott: added a wildcard *.wmflabs.org domain pointing at the domain proxy in project-proxy * 04:00 andrewbogott: shortened the ttl on .wmflabs.org. to 300 === 2020-07-13 === * 16:17 arturo: icinga downtime cloudcontrol[1003-1005].wikimedia.org for 1h for galera database movements === 2020-07-12 === * 17:39 andrewbogott: switched eqiad1 keystone from m5 to cloudcontrol galera === 2020-07-10 === * 20:26 andrewbogott: disabling nova api to move database to galera === 2020-07-09 === * 11:23 arturo: [codfw1dev] rebooting cloudnet2003-dev again for testing sysct/puppet behavior ([[phab:T257552|T257552]]) * 11:11 arturo: [codfw1dev] rebooting cloudnet2003-dev for testing sysct/puppet behavior ([[phab:T257552|T257552]]) * 09:16 arturo: manually increasing sysctl value of net.nf_conntrack_max in cloudnet servers ([[phab:T257552|T257552]]) === 2020-07-06 === * 15:16 arturo: installing 'aptitude' in all cloudvirts === 2020-07-03 === * 12:51 arturo: [codfw1dev] galera cluster should be up and running, openstack happy ([[phab:T256283|T256283]]) * 11:44 arturo: [codfw1dev] restoring glance database backup from bacula into cloudcontrol2001-dev ([[phab:T256283|T256283]]) * 11:39 arturo: [codfw1dev] stopped mysql database in the galera cluster [[phab:T256283|T256283]] * 11:36 arturo: [codfw1dev] dropped glance database in the galera cluster [[phab:T256283|T256283]] === 2020-07-02 === * 15:41 arturo: `sudo wmcs-openstack --os-compute-api-version 2.55 flavor create --private --vcpus 8 --disk 300 --ram 16384 --property aggregate_instance_extra_specs:ceph=true --description "for packaging envoy" bigdisk-ceph` ([[phab:T256983|T256983]]) === 2020-06-29 === * 14:24 arturo: starting rabbitmq-server in all 3 cloudcontrol servers * 14:23 arturo: stopping rabbitmq-server in all 3 cloudcontrol servers === 2020-06-18 === * 20:38 andrewbogott: rebooting cloudservices2003-dev due to a mysterious 'host down' alert on a secondary ip === 2020-06-16 === * 15:38 arturo: created by hand neutron port 9c0a9a13-e409-49de-9ba3-{{Gerrit|bc8ec4801dbf}} `paws-haproxy-vip` ([[phab:T295217|T295217]]) === 2020-06-12 === * 13:23 arturo: DNS zone `paws.wmcloud.org` transferred to the PAWS project ([[phab:T195217|T195217]]) * 13:20 arturo: created DNS zone `paws.wmcloud.org` ([[phab:T195217|T195217]]) === 2020-06-11 === * 19:19 bstorm_: proceeding with failback to labstore1004 now that DRBD devices are consistent [[phab:T224582|T224582]] * 17:22 bstorm_: delaying failback labstore1004 for drive syncs [[phab:T224582|T224582]] * 17:17 bstorm_: failing NFS back to labstore1004 to complete the upgrade process [[phab:T224582|T224582]] * 16:15 bstorm_: failing over NFS for labstore1004 to labstore1005 [[phab:T224582|T224582]] === 2020-06-10 === * 16:09 andrewbogott: deleting all old cloud-ns0.wikimedia.org and cloud-ns1.wikimedia.org ns records in designate database [[phab:T254496|T254496]] === 2020-06-09 === * 15:25 arturo: icinga downtime everything cloud* lab* for 2h more ([[phab:T253780|T253780]]) * 14:09 andrewbogott: stopping puppet, all designate services and all pdns services on cloudservices1004 for [[phab:T253780|T253780]] * 14:01 arturo: icinga downtime everything cloud* lab* for 2h ([[phab:T253780|T253780]]) === 2020-06-05 === * 15:08 andrewbogott: trying to re-enable puppet without losing cumin contact, as per https://phabricator.wikimedia.org/T254589 === 2020-06-04 === * 14:24 andrewbogott: disabling puppet on all instances for /labs/private recovery * 14:23 arturo: disabling puppet on all instances for /labs/private recovery === 2020-05-28 === * 23:02 bd808: `/usr/local/sbin/maintain-dbusers --debug harvest-replicas` ([[phab:T253930|T253930]]) * 13:36 andrewbogott: rebuilding cloudservices2002-dev with Buster * 00:33 andrewbogott: shutting down cloudservices2002-dev to see if we can live without it. This is in anticipation or rebuilding it entirely for [[phab:T253780|T253780]] === 2020-05-27 === * 23:29 andrewbogott: disabling the backup job on cloudbackup2001 (just like last week) so the backup doesn't start while Brooke is rebuilding labstore1004 tomorrow. * 06:03 bd808: `systemctl start mariadb` on clouddb1001 following reboot (take 2) * 05:58 bd808: `systemctl start mariadb` on clouddb1001 following reboot * 05:53 bd808: Hard reboot of clouddb1001 via Horizon. Console unresponsive. === 2020-05-25 === * 16:35 arturo: [codfw1dev] created zone `0-29.57.15.185.in-addr.arpa.` ([[phab:T247972|T247972]]) === 2020-05-21 === * 19:23 andrewbogott: disabling puppet on cloudbackup2001 to prevent the backup job from starting during maintenance * 19:16 andrewbogott: systemctl disable block_sync-tools-project.service on cloudbackup2001.codfw.wmnet to avoid stepping on current upgrade * 15:48 andrewbogott: re-imaging cloudnet1003 with Buster === 2020-05-19 === * 22:59 bd808: `apt-get install mariadb-client` on cloudcontrol1003 * 21:12 bd808: Migrating wcdo.wcdo.eqiad.wmflabs to cloudvirt1023 ([[phab:T251065|T251065]]) === 2020-05-18 === * 21:37 andrewbogott: rebuilding cloudnet2003-dev with Buster === 2020-05-15 === * 22:10 bd808: Added reedy as projectadmin in cloudinfra project ([[phab:T249774|T249774]]) * 22:05 bd808: Added reedy as projectadmin in admin project ([[phab:T249774|T249774]]) * 18:44 bstorm_: rebooting cloudvirt-wdqs1003 [[phab:T252831|T252831]] * 15:47 bd808: Manually running wmcs-novastats-dnsleaks from cloudcontrol1003 ([[phab:T252889|T252889]]) === 2020-05-14 === * 23:28 bstorm_: downtimed cloudvirt1004/6 and cloudvirt-wdqs1003 until tomorrow around this time [[phab:T252831|T252831]] * 22:21 bstorm_: upgrading qemu-system-x86 on cloudvirt1006 to backports version [[phab:T252831|T252831]] * 22:15 bstorm_: changing /etc/libvirt/qemu.conf and restarting libvirtd on cloudvirt1006 [[phab:T252831|T252831]] * 21:12 andrewbogott: rebuilding cloudvirt1003-wdqs as part of [[phab:T252831|T252831]] * 15:47 andrewbogott: moving cloudvirt1004 and cloudvirt1006 to the 'ceph' aggregate for [[phab:T252784|T252784]] * 15:02 andrewbogott: moving all of cloudvirt100[1-9] into the 'toobusy' host aggregate. These are slower, have spinning disks, and are due for replacement. === 2020-05-12 === * 20:33 andrewbogott: moving cloudvirt1023 to the 'standard' pool and out of the 'spare' pool * 19:10 jeh: disable neutron-openvswitch-agent service on cloudvirt2001-dev.codfw [[phab:T248881|T248881]] * 19:09 jeh: Shutdown the unused eno2 network interface on cloudvirt2001-dev.codfw to clear up monitoring errors [[phab:T248425|T248425]] * 18:20 andrewbogott: moving cloudvirt1024 out of the 'maintenance' aggregate and into 'spare' * 16:45 andrewbogott: restarting neutron-l3-agent on cloudnet1004 so it knows about all three cloudcontrols. Leaving cloudnet1003 since restarting it there will cause network interruptions * 14:06 arturo: icinga downtime everything for 2h for Debian Buster migration in some cloud components === 2020-05-09 === * 16:53 andrewbogott: rebuilding cloudcontrol2001-dev and 2003-dev with buster for [[phab:T252121|T252121]] === 2020-05-08 === * 19:02 bstorm_: moving tools-k8s-haproxy-2 from cloudvirt1021 to cloudvirt1017 to improve spread === 2020-05-05 === * 13:58 andrewbogott: rebuilding cloudcontrol2004-dev to test new puppet changes === 2020-05-04 === * 09:04 arturo: [codfw1dev] manually modify iptables ruleset to only allow SSH from WMF bastions on cloudservices2003-dev and cloudcontrol2004-dev ([[phab:T251604|T251604]]) === 2020-04-21 === * 22:12 andrewbogott: moving cloudvirt1004 out of the 'standard' aggregate and into the 'maintenance' aggregate * 16:01 jeh: restart cloudceph mon and osd services for openssl upgrades === 2020-04-15 === * 18:44 jeh: create indexes and views for grwikimedia [[phab:T245912|T245912]] === 2020-04-13 === * 15:07 jeh: restart memcached on labwebs to increase cache size [[phab:T145703|T145703]] === 2020-04-09 === * 19:57 andrewbogott: upgrading eqiad1 designate to rocky * 16:52 andrewbogott: cleaned up a bunch of leaked .eqiad.wmflabs dns records === 2020-04-08 === * 19:20 andrewbogott: rotated password and api token for pdns servers on cloudservices1003 and cloudservices1004 * 14:54 arturo: `root@cloudcontrol1003:~# cp /etc/inputrc .inputrc` to solve some bash shortcut weirdness === 2020-04-07 === * 20:57 andrewbogott: service sssd stop; rm -rf /var/lib/sss/db*; service sssd start on tools-sgebastion-08 === 2020-04-06 === * 22:39 andrewbogott: deleting bogus groups cn=b'project-bastion',ou=groups,dc=wikimedia,dc=org and cn=b'project-tools',ou=groups,dc=wikimedia,dc=org from ldap * 17:42 arturo: [codfw1dev] transferred DNS zone 57.15.185.in-addr.arpa. to the cloudinfra-codfw1dev project ([[phab:T247972|T247972]]) * 17:39 arturo: [codfw1dev] `openstack zone create --email root@wmflabs.org --type PRIMARY --ttl 3600 --description "floating IPs subnet" 57.15.185.in-addr.arpa.` ([[phab:T247972|T247972]]) * 16:23 arturo: restarting apache2 in cloudcontrol1003/1004 to pick up latest wmfkeystonehooks changes [[phab:T249494|T249494]] === 2020-04-02 === * 20:59 jeh: codfw1dev clear VM error states and start bastions, puppet master and database === 2020-04-01 === * 16:27 arturo: [codfw1dev] enable puppet across the fleet clean vxlan changes ([[phab:T248881|T248881]]) === 2020-03-31 === * 12:35 arturo: [codfw1dev] restarting VMs: designaterockytest14, bastion-codfw1dev-0[1,2] ([[phab:T248881|T248881]]) * 12:34 arturo: [codfw1dev] installing neutron-openvswitch-agent on cloudvirt2001-dev ([[phab:T248881|T248881]]) * 12:25 arturo: [codfw1dev] installing neutron-openvswitch-agent on cloudnet200[2,3]-dev ([[phab:T248881|T248881]]) * 11:45 arturo: [codfw1dev] rebooting cloudvirt2003-dev to pick up latest kernel update. Otherwise modprobe is confused trying to load modules and openvswitch won't start ([[phab:T248881|T248881]]) * 10:40 arturo: [codfw1dev] installing neutron-openvswitch-agent on cloudvirt2003-dev ([[phab:T248881|T248881]]) * 10:09 arturo: [codfw1dev] reboot cloudnet2003-dev into linux 4.9 (was using 4.14 from a testing operation in 2020-03-10) === 2020-03-30 === * 23:42 bstorm_: deleted "Kubernetes Cluster" and "Kubernetes Performance" dashboards [[phab:T246689|T246689]] * 16:44 arturo: [codfw1dev] installing package neutron-openvswitch-agent in cloudvirt2002-dev ([[phab:T248881|T248881]]) * 16:42 andrewbogott: restarting l3 agents on cloudnets in codfw1dev after applying https://gerrit.wikimedia.org/r/#/c/operations/puppet/+/584188/ === 2020-03-27 === * 21:28 bd808: Created huggle.wmcloud.org Designate zone and allocated it to the huggle project * 19:51 jeh: start haproxy on cloudcontrol2003-dev.wikimedia.org === 2020-03-26 === * 15:01 arturo: icinga downtime cloudvirt* cloudcontrol* cloudnet* lab* cloudstore* * 15:01 andrewbogott: beginning openstack upgrade window for [[phab:T242766|T242766]] * 12:32 arturo: [codfw1dev] downgraded systemd, libsystemd0, udev and friends to the non-backports versions ([[phab:T247013|T247013]]) === 2020-03-25 === * 19:29 andrewbogott: dumping a bunch of VMs on cloudvirt1015 to see if it still crashes * 17:56 jeh: add labweb1002 back into the pool - completed horizon testing [[phab:T240852|T240852]] * 17:09 jeh: depool labweb1002 for horizon testing [[phab:T240852|T240852]] === 2020-03-24 === * 19:41 jeh: switch cloudvirt1016 from maintenance to standard host aggregate [[phab:T243327|T243327]] * 15:31 andrewbogott: restarting nova-conductor and nova-api on cloudcontrol1003 and cloudcontrol1004 === 2020-03-23 === * 21:41 jeh: restart neutron-l3-agent on cloudnet100[3,4] to pickup policy.yaml changes * 13:28 jeh: disable puppet on labweb100[1,2] to enable horizon event traces [[phab:T240852|T240852]] * 10:26 arturo: restarting apache in both labweb1001/labweb1002 upon reports of returning 500s === 2020-03-21 === * 14:23 andrewbogott: restarting apache2 on labweb1001 and 1002 === 2020-03-18 === * 19:17 andrewbogott: deleted a bunch of records from the pdns database on cloudservices1003/1004 which had a record name but the content (where an IP address should be) was NULL, e.g. m.wikidata.beta.wmflabs.org. * 10:55 arturo: [codfw1dev] deleting BGP agent, undoing changes we did for [[phab:T245606|T245606]] === 2020-03-14 === * 17:40 jeh: restart maintain-dbusers on labstore1004 [[phab:T247654|T247654]] === 2020-03-13 === * 12:39 arturo: [codfw1dev] reintroduce address scopes for another round of testing [[phab:T244851|T244851]] * 12:17 arturo: [codfw1dev] enabling puppet in cloudnet200x-dev servers after merging https://gerrit.wikimedia.org/r/c/operations/puppet/+/579259 ([[phab:T247505|T247505]]) === 2020-03-12 === * 22:29 bstorm_: running puppet across all dumps mounts to make sure active links are shifted to labstore1006 === 2020-03-11 === * 18:38 jeh: set icingia downtime until 2020-03-23 on CODFW cloud[control,net,virt] hosts during openstack upgrades * 12:50 arturo: [codfw1dev] several tests creating/deleting address scopes ([[phab:T244727|T244727]] [[phab:T247135|T247135]] [[phab:T246887|T246887]] [[phab:T245606|T245606]]) * 12:46 arturo: [codfw1dev] disable routing_source_ip in l3 agents for testing proposal detailed at https://wikitech.wikimedia.org/wiki/Wikimedia_Cloud_Services_team/EnhancementProposals/Network_refresh#Eliminate_routing_source_ip_address ([[phab:T244727|T244727]]) === 2020-03-10 === * 17:02 arturo: [codfw1dev] deleting address scopes, bad interaction with our custom NAT setup [[phab:T247135|T247135]] * 13:55 arturo: [codfw1dev] rebooting cloudnet2003-dev into linux kernel 4.14 for testing stuff related to [[phab:T247135|T247135]] === 2020-03-09 === * 18:09 arturo: enabling puppet in cloudvirt1006, all services have been restored * 17:59 arturo: deleted the neutron bridge on cloudvirt1006, for testing stuff related to the queens upgrade * 17:58 arturo: stopped neutron-linuxbridge-agent and nova-compute in cloudvirt1006 for testing stuff related to the queens upgrade === 2020-03-06 === * 14:54 andrewbogott: draining all instances off of cloudvirt1006 for [[phab:T246908|T246908]] === 2020-03-05 === * 14:24 arturo: [codfw1dev] we just enabled BGP session between cloudnet2xxx-dev and cr1-codfw ([[phab:T245606|T245606]]) * 13:07 arturo: [codfw1dev] move the extra IP address for BGP in cloudnet200x-dev servers from eno2.2120 to the br-external bridge device ([[phab:T245606|T245606]]) * 13:06 arturo: [codfw1dev] upgrade neutron-dynamic-routing packages in cloudnet200X-dev and cloudcontrol200X-dev servers to 11.0.0-2~bpo9+1 ([[phab:T245606|T245606]]) === 2020-03-04 === * 22:22 andrewbogott: upgrading designate on cloudservices1003/1004 to Queens * 22:09 andrewbogott: moving cloudvirt1006 into the maintenance aggregate for [[phab:T246908|T246908]] * 21:37 bd808: Running wmcs-wikireplica-dns to add service names for ngwikimedia.*.db.svc.eqiad.wmflabs ([[phab:T240772|T240772]]) * 21:14 bd808: Running `sudo maintain-meta_p --all-databases --purge` on labsdb1009 ([[phab:T246056|T246056]]) * 21:11 bd808: Running `sudo maintain-meta_p --all-databases --purge` on labsdb1010 ([[phab:T246056|T246056]]) * 21:08 bd808: Running `sudo maintain-meta_p --all-databases --purge` on labsdb1011 ([[phab:T246056|T246056]]) * 21:05 bd808: Running `sudo maintain-meta_p --all-databases --purge` on labsdb1002 ([[phab:T246056|T246056]]) === 2020-03-02 === * 16:54 arturo: [codfw1dev] deleted python3-os-ken debian package in cloudnet2003-dev which was installed by hand and had depedency issues === 2020-02-29 === * 16:32 bstorm_: downtimed the smart alert on cloudvirt1009 until Monday since apparently predictive failures flap [[phab:T244986|T244986]] === 2020-02-26 === * 22:03 jeh: powering down cloudvirt1014 for hardware maintenance === 2020-02-25 === * 16:08 andrewbogott: changing neutron's rabbitmq password because oslo is having trouble parsing some of the characters in the password * 15:26 andrewbogott: updated the cell_mapping record in the nova_api database to add the second rabbitmq server to the transport_url field * 15:26 andrewbogott: updated the cell_mapping record in the nova_api database to set the db uri to 'mysql+pymysql' -- this in response to a deprecation notice === 2020-02-24 === * 12:16 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# neutron bgp-speaker-peer-add bgpspeaker cr2-codfw` ([[phab:T245606|T245606]]) * 12:16 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# neutron bgp-speaker-peer-add bgpspeaker cr1-codfw` ([[phab:T245606|T245606]]) * 12:09 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# neutron bgp-peer-create --peer-ip 208.80.153.187 --remote-as 65002 cr2-codfw` ([[phab:T245606|T245606]]) * 12:09 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# neutron bgp-peer-create --peer-ip 208.80.153.186 --remote-as 65002 cr1-codfw` ([[phab:T245606|T245606]]) * 12:06 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# neutron bgp-peer-delete 17b8c2a3-f0ce-4d50-a265-18ccac703c61` ([[phab:T245606|T245606]]) * 10:59 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# neutron bgp-speaker-peer-add bgpspeaker bgppeer` ([[phab:T245606|T245606]]) * 10:56 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# neutron bgp-peer-create --peer-ip 208.80.153.185 --remote-as 65002 bgppeer` ([[phab:T245606|T245606]]) === 2020-02-21 === * 12:48 arturo: [codfw1dev] running `root@cloudcontrol2001-dev:~# neutron bgp-speaker-network-add bgpspeaker wan-transport-codfw` ([[phab:T245606|T245606]]) * 12:46 arturo: [codfw1dev] created bgpspeaker for AS64711 ([[phab:T245606|T245606]]) * 12:42 arturo: [codfw1dev] run `sudo neutron-db-manage upgrade head` to upgrade the db schema for neutron bgp tables * 11:51 arturo: [codfw1dev] create a neutron subnet pool per each subnet objects we have and manually update DB to inter-associate them ([[phab:T245606|T245606]]) * 11:49 arturo: [codfw1dev] rename neutron address scope `no-nat` to `bgp` ([[phab:T245606|T245606]]) * 11:37 arturo: [codfw1dev] cleanup unused neutron subnet pools from previous address scope tests ([[phab:T244851|T244851]]) === 2020-02-20 === * 19:22 andrewbogott: updating designate pool config for https://gerrit.wikimedia.org/r/#/c/operations/puppet/+/572213/ * 15:33 andrewbogott: migrating all VMs on cloudvirt1014 to cloudvirt1022 * 13:35 arturo: [codfw1dev] disable puppet in cloudcontrol servers to hack neutron.conf for tests related to [[phab:T245606|T245606]] * 13:33 arturo: [codfw1dev] disable puppet in cloudnet servers to hack neutron.conf for tests related to [[phab:T245606|T245606]] === 2020-02-18 === * 22:19 andrewbogott: transferred the tools.wmcloud.org. to the tools project * 22:16 andrewbogott: moved wmcloud.org dns domain to the cloud-infra project * 21:02 andrewbogott: adding .eqiad1.wikimedia.cloud records to all existing eqiad1 VMs, updating all eqiad1 internal pointer records to reference the new eqiad1.wikimedia.cloud fqdns. * 09:44 arturo: deleted DNS zone wmcloud.org and try re-creating it === 2020-02-14 === * 10:35 arturo: running `root@cloudcontrol2001-dev:~# designate server-create --name ns1.openstack.codfw1dev.wikimediacloud.org.` ([[phab:T243766|T243766]]) * 10:32 arturo: running `root@cloudcontrol1004:~# designate server-create --name ns1.openstack.eqiad1.wikimediacloud.org.` ([[phab:T243766|T243766]]) * 10:32 arturo: running `root@cloudcontrol1004:~# designate server-create --name ns0.openstack.eqiad1.wikimediacloud.org.` ([[phab:T243766|T243766]]) === 2020-02-12 === * 13:38 arturo: [codfw1dev] add reference to subnetpool to the instance subnet `MariaDB [neutron]> update subnets set subnetpool_id='d129650d-d4be-4fe1-b13e-6edb5565cb4a' where id = '7adfcebe-b3d0-4315-92fe-e8365cc80668';` ([[phab:T244851|T244851]]) === 2020-02-11 === * 13:46 arturo: [codfw1dev] creating some neutron objects to investigate [[phab:T244851|T244851]] (subnets, subnet pools, address scopes, ...) * 12:40 arturo: [codfw1dev] delete unknown address scope 'wmcs-v4-scope': `root@cloudcontrol2001-dev:~# openstack address scope delete 078cfd71-117b-4aac-9197-6ebbbb7dd3de` ([[phab:T244851|T244851]]) * 12:40 arturo: [codfw1dev] delete unknown subnet pool 'cloudinstancesb-v4-pool0': `root@cloudcontrol2001-dev:~# openstack subnet pool delete d23a9b88-5c3d-4a53-ab88-053233a75365` ([[phab:T244851|T244851]]) === 2020-02-07 === * 18:11 jeh: shutdown cloudvirt1016 for hardware maintenance [[phab:T241882|T241882]] === 2020-02-06 === * 14:44 jeh: update apt packages on cloudvirt1015 [[phab:T220853|T220853]] * 14:28 jeh: run hardware tests on cloudvirt1015 [[phab:T220853|T220853]] === 2020-01-28 === * 17:24 arturo: [codfw1dev] root@cloudcontrol2001-dev:~# designate server-create --name ns0.openstack.codfw1dev.wikimediacloud.org. ([[phab:T243766|T243766]]) * 10:18 arturo: [codfw1dev] created DNS record `bastion-codfw1dev-01.codfw1dev.wmcloud.org A 185.15.57.2` ([[phab:T242976|T242976]], [[phab:T229441|T229441]]) * 10:13 arturo: [codfw1dev] the zone `codfw1dev.wmcloud.org` belongs now to the `cloudinfra-codfw1dev` project ([[phab:T242976|T242976]]) * 10:11 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# openstack zone create --description "main DNS domain for public addresses" --email "root@wmflabs.org" --type PRIMARY --ttl 3600 codfw1dev.wmcloud.org.` ([[phab:T242976|T242976]] and [[phab:T243766|T243766]]) * 09:53 arturo: restart apache2 in labweb1001/1002 because horizon errors * 09:47 arturo: created DNS zone wmcloud.org in eqiad1, transfer it to the cloudinfra project ([[phab:T242976|T242976]]) right now only use is to delegate codfw1dev.wmcloud.org subdomain to designate in the other deployment === 2020-01-27 === * 12:45 arturo: [codfw1dev] manually move the new domain to the `cloudinfra-codfw1dev` project clouddb2001-dev: `[designate]> update zones set tenant_id='cloudinfra-codfw1dev' where id = '4c75410017904858a5839de93c9e8b3d';` [[phab:T243556|T243556]] * 12:44 arturo: [codfw1dev] `root@cloudcontrol2001-dev:~# openstack zone create --description "main DNS domain for VMs" --email "root@wmflabs.org" --type PRIMARY --ttl 3600 codfw1dev.wikimedia.cloud.` [[phab:T243556|T243556]] === 2020-01-24 === * 15:10 jeh: remove icinga downtime for cloudvirt1013 [[phab:T241313|T241313]] * 12:52 arturo: repooling cloudvirt1013 after HW got fixed ([[phab:T241313|T241313]]) === 2020-01-21 === * 17:43 bstorm_: remounting /mnt/nfs/dumps-labstore1007.wikimedia.org/ on all dumps-mounting projects * 10:24 arturo: running `sudo systemctl restart apache2.service` in both labweb servers to try mitigating [[phab:T240852|T240852]] === 2020-01-15 === * 16:59 bd808: Changed the config for cloud-announce mailing list so that lsit admins do not get bounce unsubscribe notices === 2020-01-14 === * 14:03 arturo: icinga downtime all cloudvirts for another 2h for fixing some icinga checks * 12:04 arturo: icinga downtime toolchecker for 2 hours for openstack upgrades [[phab:T241347|T241347]] * 12:02 arturo: icinga downtime cloud* labs* hosts for 2 hours for openstack upgrades [[phab:T241347|T241347]] * 04:26 andrewbogott: upgrading designate on cloudservices1003/1004 === 2020-01-13 === * 13:34 arturo: [¢odfw1dev] prevent neutron from allocating floating IPs from the wrong subnet by doing `neutron subnet-update --allocation-pool start=208.80.153.190,end=208.80.153.190 cloud-instances-transport1-b-codfw` ([[phab:T242594|T242594]]) === 2020-01-10 === * 13:27 arturo: cloudvirt1009: virsh undefine i-000069b6. This is tools-elastic-01 which is running on cloudvirt1008 (so, leaked on cloudvirt1009) === 2020-01-09 === * 11:12 arturo: running `MariaDB [nova_eqiad1]> update quota_usages set in_use='0' where project_id='etytree';` ([[phab:T242332|T242332]]) * 11:11 arturo: running `MariaDB [nova_eqiad1]> select * from quota_usages where project_id = 'etytree';` ([[phab:T242332|T242332]]) * 10:32 arturo: ran `root@cloudcontrol1004:~# nova-manage project quota_usage_refresh --project etytree` === 2020-01-08 === * 10:53 arturo: icinga downtime all cloudvirts for 30 minutes to re-create all canary VMs" === 2020-01-07 === * 11:12 arturo: icinga-downtime everything cloud* for 30 minutes to merge nova scheduler changes * 10:02 arturo: icinga downtime cloudvirt1009 for 30 minutes to re-create canary VM ([[phab:T242078|T242078]]) === 2020-01-06 === * 13:45 andrewbogott: restarting nova-api and nova-conductor on cloudcontrol1003 and 1004 === 2020-01-04 === * 16:34 arturo: icinga downtime cloudvirt1024 for 2 months because hardware errors ([[phab:T241884|T241884]]) === 2019-12-31 === * 11:46 andrewbogott: I couldn't! * 11:40 andrewbogott: restarting cloudservices2002-dev to see if I can reproduce an issue I saw earlier === 2019-12-25 === * 10:13 arturo: icinga downtime for 30 minutes the whole cloud* lab* fleet to merge https://gerrit.wikimedia.org/r/c/operations/puppet/+/560575 (will restart some openstack components) === 2019-12-24 === * 15:13 arturo: icinga downtime all the lab* fleet for nova password change for 1h * 14:39 arturo: icinga downtime all the cloud* fleet for nova password change for 1h === 2019-12-23 === * 11:13 arturo: enable puppet in cloudcontrol1003/1004 * 10:40 arturo: disable puppet in cloudcontrol1003/1004 while doing changes related to python-ldap === 2019-12-22 === * 23:48 andrewbogott: restarting nova-conductor and nova-api on cloudcontrol1003 and 1004 * 09:45 arturo: cloudvirt1013 is back (did it alone) [[phab:T241313|T241313]] * 09:37 arturo: cloudvirt1013 is down for good. Apparently powered off. I can't even reach it via iLO === 2019-12-20 === * 12:43 arturo: icinga downtime cloudmetrics1001 for 128 hours === 2019-12-18 === * 12:55 arturo: [codfw1dev] created a new subnet neutron object to hold the new CIDR for floating IPs (cloud-codfw1dev-floating - 185.15.57.0/29) [[phab:T239347|T239347]] === 2019-12-17 === * 07:21 andrewbogott: deploying horizon/train to labweb1001/1002 === 2019-12-12 === * 06:11 arturo: schedule 4h downtime for labstores * 05:57 arturo: schedule 4h downtime for cloudvirts and other openstack components due to upgrade ops === 2019-12-02 === * 06:28 andrewbogott: running nova-manage db sync on eqiad1 * 06:27 andrewbogott: running nova-manage cell_v2 map_cell0 on eqiad1 === 2019-11-21 === * 16:07 jeh: created replica indexes and views for szywiki [[phab:T237373|T237373]] * 15:48 jeh: creating replica indexes and views for shywiktionary [[phab:T238115|T238115]] * 15:48 jeh: creating replica indexes and views for gcrwiki [[phab:T238114|T238114]] * 15:46 jeh: creating replica indexes and views for minwiktionary [[phab:T238522|T238522]] * 15:36 jeh: creating replica indexes and views for gewikimedia [[phab:T236404|T236404]] === 2019-11-18 === * 19:27 andrewbogott: repooling labsdb1011 * 18:54 andrewbogott: running maintain-views --all-databases --replace-all —clean on labsdb1011 [[phab:T238480|T238480]] * 18:44 andrewbogott: depooling labsdb1011 and killing remaining user queries [[phab:T238480|T238480]] * 18:42 andrewbogott: repooled labsdb1009 and 1010 [[phab:T238480|T238480]] * 18:19 andrewbogott: running maintain-views --all-databases --replace-all —clean on labsdb1010 [[phab:T238480|T238480]] * 18:18 andrewbogott: depooling labsdb1010, killing remaining user queries * 17:46 andrewbogott: running maintain-views --all-databases --replace-all —clean on labsdb1009 [[phab:T238480|T238480]] * 17:38 andrewbogott: depooling labsdb1009, killing remaining user queries * 16:54 andrewbogott: running maintain-views --all-databases --replace-all —clean on labsdb1012 [[phab:T237509|T237509]] === 2019-11-15 === * 20:04 andrewbogott: repool labdb1011 ([[phab:T237509|T237509]]) * 19:29 andrewbogott: running maintain-views --all-databases --replace-all —clean on labsdb1011 * 19:25 andrewbogott: depooling labsdb1011, killing remaining queries * 19:25 andrewbogott: repooling labsdb1010 * 18:59 andrewbogott: running maintain-views --all-databases --replace-all —clean on labsdb1012 * 18:57 andrewbogott: running maintain-views --all-databases --replace-all —clean on labsdb1010 * 18:54 andrewbogott: depooling labsdb1010, killing remaining user queries * 18:54 andrewbogott: depooled labsdb1009, ran maintain-views —clean —all-databases —replace-all, repooled === 2019-11-11 === * 13:10 arturo: cloudweb2001-dev: disable puppet and redirect stderr in the loadExitNodes.php cron script to prevent cronspam while we investigate the cause of the issue ([[phab:T237971|T237971]]) === 2019-11-05 === * 11:59 arturo: icinga downtime for 1h cloudcontrol1004, cloudnet1003, cloudvirt1017/1020/1022 for PDU operations in the rack [[phab:T227542|T227542]] === 2019-11-04 === * 21:55 andrewbogott: deleting a ton of wikitech hiera pages that were either no-ops or refer to nonexistent VMs or prefixes === 2019-10-31 === * 11:01 arturo: icinga-downtimed cloudvirt1030 and cloudservices1003 for 1h due to PDU upgrade operations [[phab:T227543|T227543]] === 2019-10-30 === * 22:43 jeh: reboot cloud-bootstrapvz-stretch to resolve bad bootstrapvz build === 2019-10-29 === * 10:52 arturo: icinga downtime cloudvirt1001/1002/1024/1018/1012/1009/1015/1008 for 1h [[phab:T227538|T227538]] === 2019-10-25 === * 10:45 arturo: icinga downtime toolschecker for 1 to upgrade clouddb1002 mariadb (toolsdb secondary) ([[phab:T236384|T236384]] , [[phab:T236420|T236420]]) === 2019-10-24 === * 12:30 arturo: starting cloudvirt1019, PDU operations ended ([[phab:T227540|T227540]]) * 11:58 arturo: icinga downtime for 2h ([[phab:T227540|T227540]]) cloudvirt1019 * 11:15 arturo: poweroff cloudvirt1019 during the PDU operations ([[phab:T227540|T227540]]) * 11:10 arturo: icinga downtime for 2h ([[phab:T227540|T227540]]) toolschecker * 10:58 arturo: icinga downtime for 1h ([[phab:T227540|T227540]]) cloudvirt100[3-7], cloudvirt1019, cloudvirt1016, cloudvirt1021, cloudvirt1013, cloudnet1004 === 2019-10-23 === * 09:23 arturo: cloudvirt1026 reboot ended OK * 09:12 arturo: rebooting cloudvirt1026 for kernel upgrade * 09:09 arturo: cloudvirt1025 reboot ended OK * 09:00 arturo: rebooting cloudvirt1025 for kernel upgrade * 08:51 arturo: icinga downtime cloudvirt1025/1026 for reboots === 2019-10-18 === * 16:01 arturo: created the `eqiad1.wikimedia.cloud` DNS zone ([[phab:T235846|T235846]]) * 14:27 andrewbogott: deleted a bunch of leaked VMS from earlier today from the admin-monitoring project. Fullstack leaks due to an api outage, maybe? * 10:44 arturo: double max_message_size from 40KB to 80KB in the cloud-admin mailing list. A simple email with a couple of quotes can go over the 40KB limit. === 2019-10-16 === * 21:59 jeh: resync wiki replica tool and user accounts [[phab:T235697|T235697]] * 09:40 arturo: reboot of cloudvirt1030 went fine * 09:28 arturo: reboot of cloudvirt1029 went fine * 09:28 arturo: rebooting cloudvirt1030 for kernel updates * 09:12 arturo: rebooting cloudvirt1029 for kernel updates * 09:11 arturo: reboot of cloudvirt1028 went fine * 09:00 arturo: rebooting cloudvirt1028 for kernel updates * 08:56 arturo: icinga downtime cloudvirt[1028-1030].eqiad.wmnet for 1h for reboots === 2019-10-15 === * 13:30 jeh: creating indexes and views for banwiki [[phab:T234770|T234770]] === 2019-10-10 === * 18:55 bd808: Created indexes and views for nqowiki ([[phab:T230543|T230543]]) * 11:59 arturo: network switch hardware is down affecting cloudvirt1025/1026 ([[phab:T227536|T227536]]) VMs are supposed to be online but unreachable === 2019-10-09 === * 10:44 arturo: cloudvirt1013 rebooted well * 10:32 arturo: cloudvirt1013 is rebooting * 10:32 arturo: cloudvirt1012 rebooted just fine (very slow, 35 VMs) * 10:21 arturo: cloudvirt1012 is rebooting * 10:19 arturo: cloudvirt1009 rebooted just fine (very slow though) * 10:07 arturo: cloudvirt1009 is rebooting * 10:06 arturo: cloudvirt1008 rebooted just fine (very slow though) * 09:58 arturo: cloudvirt1008 is rebooting * 09:52 arturo: icinga downtime toolschecker, paws, etc for 2h, because cloudvirt reboots === 2019-10-07 === * 14:07 arturo: horizon is disabled for maintenance ([[phab:T212302|T212302]]) * 14:00 arturo: starting scheduled maintenance: upgrading eqiad1 from openstack mitaka to newton === 2019-10-02 === * 15:23 arturo: codfw1dev renaming net/subnet objects to a more modern naming scheme [[phab:T233665|T233665]] * 12:49 arturo: codfw1dev delete all floating ip allocations in the deployment for mangling the network config for testing [[phab:T233665|T233665]] * 12:47 arturo: codfw1dev deleting all VMs in the deployment for mangling the network config for testing [[phab:T233665|T233665]] * 11:08 arturo: codfw1dev rebooting cloudnet2002-dev and cloudnet2003-dev for testing [[phab:T233665|T233665]] * 10:31 arturo: codfw1dev: add cloudinstances2b-gw router to the l3 agent in cloudnet2003-dev * 09:59 arturo: codfw1dev: cleanup leftover "HA port tenant admin" in neutron (ports from missing servers) * 09:46 arturo: codfw1dev: cleanup leftover neutron agents === 2019-09-30 === * 10:21 arturo: we installed ferm in every VM by mistake. Deleting it and forcing a puppet agent run to try to go back to a clean state. * 09:38 arturo: downtime toolschecker for 24h * 09:33 arturo: force update ferm cloud-wide (in all VMs) for [[phab:T153468|T153468]] === 2019-08-18 === * 10:39 arturo: rebooting cloudvirt1023 for new interface names configuration * 10:34 arturo: downtimed cloudvirt1023 for 2 days === 2019-08-05 === * 17:17 bd808: Set downtime on gridengine and kubernetes webservice checks in icinga until 2019-09-02 (flaky tests) === 2019-07-29 === * 20:14 bd808: Restarted maintain-kubeusers on tools-k8s-master-01 ([[phab:T194859|T194859]]) === 2019-07-25 === * 12:32 arturo: eqiad1/glance: debian-9.9-stretch image deprecates debian-9.8-stretch ([[phab:T228983|T228983]]) * 09:59 arturo: (codfw1dev) drop missing glance images ([[phab:T228972|T228972]]) * 09:32 arturo: (codfw1dev) deleting a bunch of VMs that were running in now missing hypervisors * 09:31 arturo: (codfw1dev) deleting a bunch of VMs in ERROR and SHUTDOWN state * 09:27 arturo: last log entry refers to the codfw1dev deployment * 09:27 arturo: cleanup `nova service-list` from old hypervisors (labtest*) * 09:23 arturo: refreshed nova DB grants in clouddb2001-dev for the codfw1dev deployment * 08:47 arturo: cleanup the cloud-announce pending emails (spam) === 2019-07-23 === * 19:43 andrewbogott: restarting rabbitmq-server on cloudcontrol1003 and 1004 === 2019-07-22 === * 23:44 bd808: Restarted maintain-kubeusers on tools-k8s-master-01 ([[phab:T228529|T228529]]) === 2019-07-11 === * 22:07 bd808: Ran `sudo systemctl stop designate_floating_ip_ptr_records_updater.service` on cloudcontrol1003 * 22:01 bd808: `sudo apt-get install python2.7-dbg` on cloudcontrol1003 to debug hung python process * 21:48 bd808: Ran `sudo systemctl stop designate_floating_ip_ptr_records_updater.service` on cloudcontrol1004 === 2019-06-25 === * 16:05 bstorm_: updated python3.4 to update4 wherever it was installed on Jessie VMs to prevent issues with broken update3. * 14:56 bstorm_: Updated python 3.4 on the labs-puppetmaster server === 2019-06-03 === * 15:55 arturo: [[phab:T221769|T221769]] rebooting cloudservices1003 after bootstrapping is apparently completed === 2019-05-28 === * 21:42 bstorm_: unmounting labstore1003-scratch on all cloud clients * 18:14 bstorm_: [[phab:T209527|T209527]] switched mounts from labstore1003 to cloudstore1008 for scratch === 2019-05-20 === * 17:25 arturo: [[phab:T223923|T223923]] dropped compat-network config from /etc/network/interfaces in eqiad1/codfw1dev neutron nodes * 17:22 arturo: [[phab:T223923|T223923]] dropped br-compat bridges and vlan interfaces (1102 and 2102) in eqiad1/codfw1dev neutron nodes * 17:07 arturo: [[phab:T223923|T223923]] dropped compat-network configuration from the neutron database in eqiad1 * 16:55 arturo: [[phab:T223923|T223923]] dropped compat-network configuration from the neutron database in codfw1dev === 2019-05-15 === * 17:00 andrewbogott: touching /root/firstboot_done on all VMs that cumin can reach. This will prevent firstboot.sh from running a second time if/when any of these are rebooted. [[phab:T223370|T223370]] === 2019-04-26 === * 15:51 arturo: andrew updated dns servers for the cloud-instances2-b-eqiad subnet in neutron: 208.80.154.143 and 208.80.154.24 === 2019-04-25 === * 11:14 arturo: [[phab:T221760|T221760]] increased size of conntrack table === 2019-04-24 === * 12:54 arturo: [[phab:T220051|T220051]] puppet broken in every VM in Cloud VPS, fixing right now === 2019-04-22 === * 11:14 arturo: create by hand /var/cache/labsaliaser/labs-ip-aliases.json in cloudservices2002-dev ([[phab:T218575|T218575]]) === 2019-04-16 === * 22:55 bd808: cloudcontrol2003-dev: added `exit 0` to /etc/cron.hourly/keystone to stop cron spam on partially configured cluster * 12:08 arturo: rebooting cloudvirt200[123]-dev because deep changes in config * 11:27 arturo: [[phab:T219626|T219626]] add DB grants for neutron and glnace to clouddb2001-dev (codfw1dev) * 10:37 arturo: [[phab:T219626|T219626]] replace 208.80.153.75 with 208.80.153.59 in the clouddb2001-dev database (codfw1dev deployment) * 10:30 arturo: [[phab:T219626|T219626]] replace labtestcontrol2003 with cloudcontrol2001-dev in the clouddb2001-dev database (codfw1dev deployment) === 2019-04-15 === * 13:08 arturo: [[phab:T219626|T219626]] add DB grants for keystone/nova/nova_api to clouddb2001-dev (codfw1dev) === 2019-04-13 === * 18:25 bd808: Restarted nova-compute service on cloudvirt1015 ([[phab:T220853|T220853]]) === 2019-04-11 === * 12:00 arturo: [[phab:T151704|T151704]] deploying oidentd to cloudnet1xxx servers === 2019-04-02 === * 19:52 andrewbogott: installed new base Stretch image. Updated packages, and runs apt-get dist-upgrade on first boot. === 2019-03-29 === * 14:34 andrewbogott: moving tools-static.wmflabs.org to point to tools-static-13 in eqiad1-r * 00:00 bstorm_: [[phab:T193264|T193264]] Added osm.db.svc.eqiad.wmflabs to cloud DNS === 2019-03-25 === * 00:40 bd808: Restarted maintain-dbusers on labstore1004. Process hung up on failed LDAP connection. === 2019-03-21 === * 19:32 andrewbogott: restarting keystone on cloudcontrol1003 === 2019-03-15 === * 16:00 gtirloni: increased nscd cache size ([[phab:T217280|T217280]]) === 2019-03-14 === * 19:04 gtirloni: bstorm started nfsd on labstore1006 ([[phab:T218341|T218341]]) * 16:42 gtirloni: published new debian-9.8 image ([[phab:T218314|T218314]]) === 2019-03-04 === * 19:37 bstorm_: umounted /mnt/nfs/dumps-labstore1006.wikimedia.org across all VPS projects for [[phab:T217473|T217473]] === 2019-02-26 === * 12:46 gtirloni: shutdown toolsbeta-sgegrid-master (cronspam) === 2019-02-25 === * 10:32 gtirloni: restarted nfsd on labstore1004 === 2019-02-21 === * 09:09 gtirloni: restarted uwsgi-labspuppetbackend.service on labpuppetmaster1001 * 07:42 gtirloni: created project cloudstore * 07:36 gtirloni: deleted wmcs-nfs project === 2019-02-20 === * 21:58 andrewbogott: silencing shinken and disabling puppet on shinken-02 for now === 2019-02-19 === * 12:00 gtirloni: added nagios@icinga2001.wikimedia.org to cloud-admin-feed@ allowed senders === 2019-02-18 === * 20:21 gtirloni: downtimed cloudvirt1020 * 20:12 gtirloni: ran `labs-ip-alias-dump.py` on cloudservices/labservices servers === 2019-02-15 === * 13:10 arturo: [[phab:T216239|T216239]] labvirt1019 has been drained * 12:22 arturo: [[phab:T216239|T216239]] draining labvirt1009 with a command like this: `root@cloudcontrol1004:~# wmcs-cold-migrate --region eqiad --nova-db nova 2c0cf363-c7c3-42ad-94bd-{{Gerrit|e586f2492321}} labvirt1001` * 12:02 arturo: more nova service cleanups in the database (labvirts that were reallocated to eqiad1) * 11:34 arturo: [[phab:T216190|T216190]] cleanup from nova database `nova service-delete 35` * 03:50 andrewbogott: updated VPS base images for Jessie and Stretch, now featuring Stretch 9.7 === 2019-02-11 === * 18:13 gtirloni: cleaned old metrics data in labmon1001 [[phab:T215417|T215417]] * 15:28 gtirloni: running `maintain-views --all-databases --replace-all` on labsdb1011 * 14:18 gtirloni: running `maintain-views --all-databases --replace-all` on labsdb1010 === 2019-02-08 === * 14:56 gtirloni: running `maintain-views --all-databases --replace-all` on labsdb1009 === 2019-02-06 === * 11:47 gtirloni: downtimed labmon100{1,2} [[phab:T215399|T215399]] * 00:17 bstorm_: [[phab:T214106|T214106]] deleted bstorm-test2 project to clean up === 2019-02-05 === * 10:48 arturo: labmon1001 is now part of the 'eqiad1-r' region === 2019-02-01 === * 09:54 arturo: moving canary1015-01 VM instance from cloudvirt1024 back to cloudvirt1015 === 2019-01-31 === * 12:44 arturo: [[phab:T215012|T215012]] depooling cloudvirt1015 and migrating all VMs to cloudvirt1024 === 2019-01-25 === * 20:11 gtirloni: deleted project yandex-proxy [[phab:T212306|T212306]] * 20:11 gtirloni: deleted project [[phab:T212306|T212306]] === 2019-01-24 === * 11:50 arturo: [[phab:T213925|T213925]] modify subnet cloud-instances-transport1-b-eqiad1 to avoid floating IP allocations from here * 11:07 arturo: [[phab:T214299|T214299]] failover cloudnet1003 to cloudnet1004 * 10:03 arturo: [[phab:T214299|T214299]] reimage cloudnet1004 to debian stretch * 09:51 arturo: [[phab:T214299|T214299]] failover cloudnet1004 to cloudnet1003 === 2019-01-22 === * 19:19 arturo: [[phab:T214299|T214299]] stretch cloudnet1003 is apparently all set * 18:40 arturo: [[phab:T214299|T214299]] manually delete from neutron agents from cloudnet1003 (must be added again after reimage, with new uuids) * 18:37 arturo: [[phab:T214299|T214299]] reimaging cloudnet1003 as debian stretch * 17:35 jbond42: starting roll out of apt package updates to * 14:41 gtirloni: [[phab:T214369|T214369]] deployed new jessie and stretch VM images === 2019-01-21 === * 18:29 gtirloni: installed libguestfs-tools on cloudvirt1021 === 2019-01-16 === * 14:21 andrewbogott: stopping old VPS proxies in eqiad — [[phab:T213540|T213540]] === 2019-01-15 === * 14:20 andrewbogott: changing tools.wmflabs.org to point to tools-proxy-03 in eqiad1 === 2019-01-13 === * 20:00 andrewbogott: VPS proxies are now running in eqiad1 on proxy-01. Old VMs will wait a bit for deletion. [[phab:T213540|T213540]] * 19:12 andrewbogott: moving the VPS proxy API backend to proxy-01.project-proxy.eqiad.wmflabs, as per [[phab:T213540|T213540]] * 17:11 andrewbogott: moving all VPS dynamic proxies to proxy-eqiad1.wmflabs.org aka proxy-01.project-proxy.eqiad.wmflabs, as per [[phab:T213540|T213540]] === 2019-01-09 === * 22:21 bd808: neutron quota-update --tenant-id tools --port 256 === 2019-01-08 === * 18:59 bd808: Definately did NOT delete uid=novaadmin,ou=people,dc=wikimedia,dc=org * 18:59 bd808: Deleted LDAP user uid=neutron,ou=people,dc=wikimedia,dc=org * 18:58 bd808: Deleted LDAP user uid=novaadmin,ou=people,dc=wikimedia,dc=org === 2019-01-06 === * 22:03 bd808: Set floatingip quota of 60 for tools project in eqiad1-r region ([[phab:T212360|T212360]]) === 2018-12-20 === * 17:10 arturo: [[phab:T207663|T207663]] renumbered transport network in eqiad1 === 2018-12-05 === * 17:59 arturo: [[phab:T207663|T207663]] changed labtestn transport network addressing from private to public === 2018-12-03 === * 13:25 arturo: [[phab:T202886|T202886]] create again PTR records after dnsleak.py fix === 2018-11-30 === * 14:08 arturo: running dns leaks cleanup `root@cloudcontrol1003:~# /root/novastats/dnsleaks.py --delete` === 2018-11-28 === * 17:33 gtirloni: deleted contintcloud project ([[phab:T209644|T209644]]) === 2018-11-27 === * 13:32 gtirloni: enabled DRBD stats collection on labstore100[4-5] [[phab:T208446|T208446]] === 2018-11-22 === * 07:12 gtirloni: deployed new debian-9.6-stretch image === 2018-11-21 === * 10:48 arturo: re-created compat-net as not shared in labtestn to test stuff related to [[phab:T209954|T209954]] === 2018-11-16 === * 12:43 gtirloni: armed keyholder on labpuppetmaster1001/1002 after reboots * 12:08 gtirloni: rebooted labpuppetmaster1001 ([[phab:T207377|T207377]]) * 11:57 gtirloni: rebooted labpuppetmaster1002 ([[phab:T207377|T207377]]) === 2018-11-14 === * 17:19 gtirloni: added cloudvirt1016 to scheduler pool ([[phab:T209426|T209426]]) * 15:41 gtirloni: reimaging labvirt1016 as cloudvirt1016 * 15:14 gtirloni: reset-failed systemd unit nova-scheduler on cloudcontrol1004 * 13:52 gtirloni: rebooted labservices1002 after package upgrades ([[phab:T207377|T207377]]) * 13:23 gtirloni: rebooted labstore2004 after package upgrades ([[phab:T207377|T207377]]) * 13:20 gtirloni: rebooted labstore2003 after package upgrades ([[phab:T207377|T207377]]) * 13:20 gtirloni: rebooted labstore2001/labstore2003 after package upgrades ([[phab:T207377|T207377]]) * 12:08 gtirloni: rebooted labnet1002 after package upgrades * 12:01 gtirloni: rebooted labmon1002 after package upgrades * 11:41 gtirloni: rebooted labcontrol1002 after package upgrades * 11:15 gtirloni: rebooted cloudcontrol1004 after package upgrades === 2018-11-09 === * 18:17 gtirloni: restarted neutron-linuxbridge-agent on cloudvirt1018/1023 === 2018-11-08 === * 11:00 gtirloni: Added novaproxy-02 to $CACHES * 10:50 gtirloni: Added cloudvirt1017 to eqiad1 region === 2018-11-07 === * 13:49 arturo: [[phab:T208733|T208733]] moving labvirt1017 from main deployment to eqiad1 and renaming it to cloudvirt1017 === 2018-10-22 === * 16:24 arturo: [[phab:T206261|T206261]] another update to dmz_cidr in eqiad1 * 10:26 arturo: change again in dmz_cidr in eqiad1: VMs will connect between them without NAT even when using floating IPs ([[phab:T206261|T206261]]) === 2018-10-19 === * 12:02 arturo: revert change in dmz_cidr in eqiad1 for now ([[phab:T206261|T206261]]) * 11:16 arturo: change in dmz_cidr in eqiad1: VMs will connect between them without NAT even when using floating IPs ([[phab:T206261|T206261]]) * 10:14 arturo: we have new virt servers in the eqiad1 deployment since past week and this week: cloudvirt1018, cloudvirt1023, cloudvirt1024 === 2018-09-26 === * 10:40 arturo: [[phab:T205524|T205524]] all sorts of restarts in all neutron daemons * 10:20 arturo: [[phab:T205524|T205524]] stop/start all neutron agents in cloudnet1003.eqiad.wmnet * 10:13 arturo: [[phab:T205524|T205524]] restart all agents in cloudnet1004.eqiad.wmnet * 10:10 arturo: restart neutron-server in cloudcontrol1003, investigating [[phab:T205524|T205524]] === 2018-09-24 === * 10:57 arturo: try to increase floating ip allocation pool in eqiad1. Of 185.15.56.0/25 we are using only 185.15.56.10-185.15.56.31, I don't know why. Let's use 185.15.56.2-185.15.56.126 === 2018-09-21 === * 17:18 bd808: Running `sudo maintain-meta_p --all-databases --purge` across labsdb10(09{{!}}10{{!}}11) for [[phab:T201890|T201890]] === 2018-09-17 === * 22:08 bd808: Granted gtirloni project roles of admin, projectadmin, and user === 2018-09-12 === * 11:20 arturo: [[phab:T202636|T202636]] distributing default routes using classless-static-route for all VMs in main/labtest (dnsmasq/nova-network) === 2018-09-11 === * 16:52 arturo: again, restarted nova-network after killing all dnsmasq procs in labnet1001 for [[phab:T202636|T202636]] * 16:08 arturo: restarted nova-network after killing all dnsmasq procs in labnet1001 for [[phab:T202636|T202636]] * 10:53 arturo: [[phab:T202636|T202636]] creating all the compat-network configuration in neutron * 10:36 arturo: [[phab:T202636|T202636]] creating br-compat bridge in eqiad1 for the compat network * 10:33 arturo: [[phab:T202636|T202636]] manually reserve 10.68.23.253 (in nova-network) === 2018-09-10 === * 22:46 andrewbogott: deleting all VMs on labvirt1019 and 1020 as prep for [[phab:T204003|T204003]] === 2018-08-30 === * 15:46 andrewbogott: restarting rabbitmq-server on cloudcontrol1003 * 13:07 arturo: [[phab:T202636|T202636]] internal network routing now exists in labtest/labtestn for VM to communicate with each other === 2018-08-28 === * 11:04 arturo: [[phab:T202549|T202549]] eqiad1 databases are all now running in m5-master. Mysql has been cleaned from cloudcontrol100[3,4] === 2018-08-23 === * 16:17 arturo: [[phab:T188589|T188589]] bstorm_ merged patch to reduce nova DB connection usage * 13:15 arturo: [[phab:T202115|T202115]] `root@cloudcontrol1003:~# neutron subnet-update --allocation-pool start=10.64.22.4,end=10.64.22.4 e4fb2771-a361-4add-ac4e-280cc300c59f` * 13:10 arturo: [[phab:T202115|T202115]] (was `{"start": "10.64.22.2", "end": "10.64.22.254"}` ) * 13:08 arturo: [[phab:T202115|T202115]] `root@cloudcontrol1003:~# neutron subnet-update --allocation-pool start=10.64.22.254,end=10.64.22.254 e4fb2771-a361-4add-ac4e-280cc300c59f` === 2018-08-22 === * 15:28 arturo: cleanup local glance,keystone databases in cloudcontrol1003.wikimedia.org (already in m5-master) * 15:27 arturo: cleanup local keystone database in cloudcontrol1003.wikimedia.org (already in m5-master) === 2018-08-21 === * 15:39 andrewbogott: initial test message * 10:31 arturo: eqiad1 remove leftover port for HA on labnet1004 * 10:15 arturo: test === 2018-05-07 === * 18:07 bstorm_: stopped the toolhistory job because it is totally broken and fills /tmp. === 2018-02-09 === * 00:55 bd808: Added Arturo Borrero Gonzalez and Bstorm as project members * 00:54 bd808: Removed Yuvipanda at user request ([[phab:T186289|T186289]]) {{SAL|Project Name=admin}} <noinclude>[[Category:SAL]]</noinclude> qkfin7f34e5ohz02br4u4oeeobxvtql Map of database maintenance 0 449160 2445264 2445245 2026-08-10T00:00:08Z Dexbot 30554 Bot: Updating the report 2445264 wikitext text/x-wiki {{/Header}} == Today (2026-08-10) == == Yesterday (2026-08-09) == == Last seven days == {| class="wikitable" |+ eqiad |- ! Section !! Work |- | x1 || [[phab:T433990|Optimize echo tables in x1 (T433990)]] (ladsgroup) |- |} {| class="wikitable" |+ codfw |- ! Section !! Work |- | x1 || [[phab:T433990|Optimize echo tables in x1 (T433990)]] (ladsgroup) |- |} [[Category:MariaDB]] e2it1mzdbe2yed68ff7dr84b3o4p59y Tool:Phab-ban/Log 116 453426 2445265 2443384 2026-08-10T01:48:59Z Phabbanbot 37210 Adminbdso was disabled by JJMC89 2445265 wikitext text/x-wiki <noinclude>'''Audit log of bans''' made via https://phab-ban.toolforge.org. Some bans made prior to 2023-09-01 were manually logged at [[phab:T200856]]. __NOTOC____NOINDEX__</noinclude> === 2026-08-10 === * 01:48 [[phab:p/Adminbdso|Adminbdso]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2026-08-03 === * 08:02 [[phab:p/Jonas2356|Jonas2356]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2026-07-13 === * 13:01 [[phab:p/Camrensssss|Camrensssss]] was disabled by [[phab:p/Mainframe98/|Mainframe98]] * 09:11 [[phab:p/VersatileDove231|VersatileDove231]] was disabled by [[phab:p/Lucas_Werkmeister_WMDE/|Lucas_Werkmeister_WMDE]] === 2026-07-09 === * 00:26 [[phab:p/VVMFOfffce|VVMFOfffce]] was disabled by [[phab:p/SomeRandomDeveloper/|SomeRandomDeveloper]] === 2026-06-30 === * 17:21 [[phab:p/Tomasz_Bladyniec|Tomasz_Bladyniec]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2026-06-29 === * 06:34 [[phab:p/LuniZunie|LuniZunie]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2026-05-25 === * 14:40 [[phab:p/Lysdexia|Lysdexia]] was disabled by [[phab:p/HakanIST/|HakanIST]] === 2026-05-24 === * 17:23 [[phab:p/Nawaf2296|Nawaf2296]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] === 2026-05-17 === * 01:27 [[phab:p/Chicken.Tender.331|Chicken.Tender.331]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2026-05-07 === * 05:44 [[phab:p/Gabor_Kiss_WMSE|Gabor_Kiss_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] === 2026-04-29 === * 22:31 [[phab:p/jhsoby-WMNO|jhsoby-WMNO]] was disabled by [[phab:p/Zabe/|Zabe]] * 20:37 [[phab:p/Hr574380|Hr574380]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2026-04-22 === * 18:16 [[phab:p/Datronmcka|Datronmcka]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2026-04-14 === * 18:53 [[phab:p/Apokrif|Apokrif]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2026-04-07 === * 11:18 [[phab:p/Biof|Biof]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2026-03-15 === * 11:47 [[phab:p/Bucheon606|Bucheon606]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 11:47 [[phab:p/Bucheon606|Bucheon606]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 11:29 [[phab:p/Seoulbucheonincheon|Seoulbucheonincheon]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 11:10 [[phab:p/Bucheonstation|Bucheonstation]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 11:04 [[phab:p/WMFOfffce|WMFOfffce]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 11:03 [[phab:p/WMFOfffce|WMFOfffce]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 11:03 [[phab:p/WMFOfffce|WMFOfffce]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 10:57 [[phab:p/WMFOfflce|WMFOfflce]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 10:29 [[phab:p/BucheonFac|BucheonFac]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 09:48 [[phab:p/BucheonWest|BucheonWest]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:43 [[phab:p/TheBucheon|TheBucheon]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:40 [[phab:p/BucheonIncheon|BucheonIncheon]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 09:33 [[phab:p/SkottishFinnishRadist|SkottishFinnishRadist]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:33 [[phab:p/SkottishFinnishRadist|SkottishFinnishRadist]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 09:25 [[phab:p/BucheonCityHall6|BucheonCityHall6]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:16 [[phab:p/Primefac1|Primefac1]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 09:07 [[phab:p/ScottishFimishRadish|ScottishFimishRadish]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 06:35 [[phab:p/Kgarcia181|Kgarcia181]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 04:27 [[phab:p/SldrF|SldrF]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 04:17 [[phab:p/PrimePac|PrimePac]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] * 04:01 [[phab:p/260315t1244|260315t1244]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 03:34 [[phab:p/Bucheon543|Bucheon543]] was disabled by [[phab:p/DLynch/|DLynch]] * 03:25 [[phab:p/LAG|LAG]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] * 03:13 [[phab:p/Wonmidong|Wonmidong]] was disabled by [[phab:p/Novem_Linguae/|Novem_Linguae]] === 2026-03-14 === * 08:10 [[phab:p/Bucheon2026|Bucheon2026]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 07:52 [[phab:p/BucheonCityHall|BucheonCityHall]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] * 07:14 [[phab:p/BucheonFesta|BucheonFesta]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2026-03-08 === * 14:34 [[phab:p/Unicord|Unicord]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] === 2026-02-22 === * 10:39 [[phab:p/Kredionecsresmi|Kredionecsresmi]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2026-02-15 === * 03:34 [[phab:p/alfredbeck|alfredbeck]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2026-02-11 === * 14:34 [[phab:p/Nguyentrongphu|Nguyentrongphu]] was disabled by [[phab:p/Superpes15/|Superpes15]] * 14:32 [[phab:p/Slowking4|Slowking4]] was disabled by [[phab:p/Superpes15/|Superpes15]] === 2026-02-09 === * 22:45 [[phab:p/UNIX-QUANTUM-UNIBANK-FICSIT-NETWORKS|UNIX-QUANTUM-UNIBANK-FICSIT-NETWORKS]] was disabled by [[phab:p/Zabe/|Zabe]] === 2026-01-23 === * 12:09 [[phab:p/Aboodhassanio|Aboodhassanio]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2026-01-16 === * 16:08 [[phab:p/Kimlien316|Kimlien316]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] * 12:01 [[phab:p/Batiste67400|Batiste67400]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2026-01-12 === * 22:04 [[phab:p/Claudioluna6|Claudioluna6]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2026-01-03 === * 14:31 [[phab:p/Cw95hh9|Cw95hh9]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-12-27 === * 06:28 [[phab:p/LBLaiSiNanHai|LBLaiSiNanHai]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-12-24 === * 21:36 [[phab:p/ItsLido|ItsLido]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-12-11 === * 15:50 [[phab:p/DanielJohnso|DanielJohnso]] was disabled by [[phab:p/Marostegui/|Marostegui]] === 2025-12-01 === * 20:36 [[phab:p/Lkcl|Lkcl]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-11-30 === * 03:56 [[phab:p/BscottAPL33|BscottAPL33]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-11-28 === * 21:21 [[phab:p/3894djfj.10439djf|3894djfj.10439djf]] was disabled by [[phab:p/JJMC89/|JJMC89]] * 09:35 [[phab:p/ElinagittHuB|ElinagittHuB]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-11-22 === * 22:09 [[phab:p/Aadvertising|Aadvertising]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 22:09 [[phab:p/TweakFind|TweakFind]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-11-20 === * 04:46 [[phab:p/SydneyRug|SydneyRug]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-11-14 === * 08:52 [[phab:p/Lougrammfoundation|Lougrammfoundation]] was disabled by [[phab:p/Marostegui/|Marostegui]] === 2025-11-12 === * 06:59 [[phab:p/Cesar|Cesar]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-11-05 === * 23:03 [[phab:p/Nehtechnine|Nehtechnine]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-11-02 === * 06:33 [[phab:p/Chhoundevid|Chhoundevid]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-11-01 === * 07:11 [[phab:p/TowfiqSir|TowfiqSir]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-10-14 === * 12:30 [[phab:p/Amrok84|Amrok84]] was disabled by [[phab:p/Lucas_Werkmeister_WMDE/|Lucas_Werkmeister_WMDE]] === 2025-09-24 === * 13:31 [[phab:p/100592|100592]] was disabled by [[phab:p/A_smart_kitten/|A_smart_kitten]] === 2025-09-22 === * 16:20 [[phab:p/DanishAhmedKm|DanishAhmedKm]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-09-12 === * 14:08 [[phab:p/GrimGwTK|GrimGwTK]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-09-04 === * 15:14 [[phab:p/Amarvip|Amarvip]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-08-29 === * 00:09 [[phab:p/IbJo|IbJo]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-08-27 === * 02:31 [[phab:p/W11228|W11228]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-08-25 === * 17:34 [[phab:p/Faster_than_Thunder|Faster_than_Thunder]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-08-22 === * 22:07 [[phab:p/T200856-01|T200856-01]] was disabled by [[phab:p/bd808/|bd808]] === 2025-08-19 === * 04:58 [[phab:p/tiffatk|tiffatk]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-08-15 === * 05:00 [[phab:p/Totrue89|Totrue89]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-08-06 === * 22:20 [[phab:p/Sajidali110|Sajidali110]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-07-31 === * 10:17 [[phab:p/EMIZENTECH|EMIZENTECH]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-07-30 === * 10:21 [[phab:p/Fariha_Asghar785|Fariha_Asghar785]] was disabled by [[phab:p/Lucas_Werkmeister_WMDE/|Lucas_Werkmeister_WMDE]] === 2025-07-08 === * 10:46 [[phab:p/Tulsi_Bhagat|Tulsi_Bhagat]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-06-29 === * 16:12 [[phab:p/sassybritches|sassybritches]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-06-22 === * 07:54 [[phab:p/Godspowertechnical|Godspowertechnical]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-06-08 === * 14:59 [[phab:p/Jaypam001|Jaypam001]] was disabled by [[phab:p/JJMC89/|JJMC89]] * 14:53 [[phab:p/DANISHAHMED111|DANISHAHMED111]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-06-07 === * 05:50 [[phab:p/PCJND|PCJND]] was disabled by [[phab:p/Johannnes89/|Johannnes89]] === 2025-06-04 === * 08:37 [[phab:p/Alpasli|Alpasli]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-06-03 === * 02:07 [[phab:p/Jj881|Jj881]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-05-29 === * 05:51 [[phab:p/RodneyAraujo|RodneyAraujo]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-04-28 === * 15:07 [[phab:p/Hansmuller|Hansmuller]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-04-03 === * 15:57 [[phab:p/Wfan|Wfan]] was disabled by [[phab:p/Zabe/|Zabe]] === 2025-03-30 === * 10:15 [[phab:p/Watnoii24|Watnoii24]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-03-23 === * 11:27 [[phab:p/Saadtbli|Saadtbli]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-03-22 === * 16:45 [[phab:p/Stephonjeffries19|Stephonjeffries19]] was disabled by [[phab:p/LucasWerkmeister/|LucasWerkmeister]] * 04:32 [[phab:p/Chriswarriortv|Chriswarriortv]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-03-19 === * 11:29 [[phab:p/Vinay080|Vinay080]] was disabled by [[phab:p/zeljkofilipin/|zeljkofilipin]] === 2025-03-18 === * 12:04 [[phab:p/Walshandpartners777|Walshandpartners777]] was disabled by [[phab:p/Lucas_Werkmeister_WMDE/|Lucas_Werkmeister_WMDE]] === 2025-03-04 === * 01:33 [[phab:p/Porokhov|Porokhov]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-02-25 === * 17:27 [[phab:p/Selahaddin751|Selahaddin751]] was disabled by [[phab:p/brennen/|brennen]] === 2025-02-19 === * 01:00 [[phab:p/Mrb_Rafi|Mrb_Rafi]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2025-02-14 === * 19:19 [[phab:p/3652candy|3652candy]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 17:01 [[phab:p/Ataysaa|Ataysaa]] was disabled by [[phab:p/bd808/|bd808]] === 2025-02-09 === * 09:10 [[phab:p/BTullis|BTullis]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2025-02-08 === * 23:26 [[phab:p/Alexdivkovic05|Alexdivkovic05]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-02-06 === * 06:19 [[phab:p/HormigasAIS|HormigasAIS]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-01-29 === * 07:40 [[phab:p/Denker61|Denker61]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2025-01-25 === * 21:36 [[phab:p/Khnthichith|Khnthichith]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-01-24 === * 11:33 [[phab:p/Aek191010|Aek191010]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2025-01-05 === * 15:40 [[phab:p/szsuperzuper|szsuperzuper]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2025-01-01 === * 09:08 [[phab:p/GALAXYENTERPRISES|GALAXYENTERPRISES]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-12-20 === * 00:52 [[phab:p/Mail.faluzes|Mail.faluzes]] was disabled by [[phab:p/Reedy/|Reedy]] === 2024-12-13 === * 02:01 [[phab:p/Gussdafii|Gussdafii]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-12-11 === * 08:54 [[phab:p/CodeTrailblazer|CodeTrailblazer]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 08:54 [[phab:p/SelvikIN|SelvikIN]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-12-03 === * 05:16 [[phab:p/Matkospajdr|Matkospajdr]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-12-01 === * 19:27 [[phab:p/Adarshsingh|Adarshsingh]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-11-28 === * 22:47 [[phab:p/Sandraklemma|Sandraklemma]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-11-23 === * 09:00 [[phab:p/Mahimabajpayee12|Mahimabajpayee12]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-11-11 === * 11:09 [[phab:p/Mvwservices|Mvwservices]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 07:00 [[phab:p/Impactolog|Impactolog]] was disabled by [[phab:p/revi/|revi]] === 2024-10-30 === * 09:05 [[phab:p/Jweighed1|Jweighed1]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-10-25 === * 04:20 [[phab:p/Blunt2531|Blunt2531]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-10-08 === * 08:31 [[phab:p/Surfcityrecovery|Surfcityrecovery]] was disabled by [[phab:p/MoritzMuehlenhoff/|MoritzMuehlenhoff]] === 2024-10-01 === * 21:49 [[phab:p/T200856-01|T200856-01]] was disabled by [[phab:p/bd808/|bd808]] === 2024-09-27 === * 10:20 [[phab:p/SorBP|SorBP]] was disabled by [[phab:p/TheresNoTime/|TheresNoTime]] === 2024-09-08 === * 10:45 [[phab:p/Robin_Mathew_Rajan|Robin_Mathew_Rajan]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2024-09-02 === * 17:58 [[phab:p/Idxntcx|Idxntcx]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-09-01 === * 10:11 [[phab:p/LDAP|LDAP]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-07-22 === * 08:56 [[phab:p/Nobleadele|Nobleadele]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-06-18 === * 20:54 [[phab:p/Playgiirlkaybrazy|Playgiirlkaybrazy]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-06-08 === * 22:30 [[phab:p/Exposingsesion1|Exposingsesion1]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-05-27 === * 12:24 [[phab:p/JosefineHellrothLarssonWMSE|JosefineHellrothLarssonWMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 07:54 [[phab:p/SMMpanels|SMMpanels]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-05-22 === * 11:31 [[phab:p/Sandra_Fauconnier_WMSE|Sandra_Fauconnier_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:23 [[phab:p/MiaJacobssonWMSE|MiaJacobssonWMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:20 [[phab:p/David_Haskiya_WMSE|David_Haskiya_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:20 [[phab:p/kalle|kalle]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:20 [[phab:p/Tore_Danielsson_WMSE|Tore_Danielsson_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:19 [[phab:p/Gitta|Gitta]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:19 [[phab:p/annatroberg|annatroberg]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:19 [[phab:p/AxelPettersson_WMSE|AxelPettersson_WMSE]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] * 10:17 [[phab:p/SaraMortsell|SaraMortsell]] was disabled by [[phab:p/Sebastian_Berlin-WMSE/|Sebastian_Berlin-WMSE]] === 2024-05-13 === * 14:55 [[phab:p/BenoitPrieur|BenoitPrieur]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2024-05-04 === * 19:54 [[phab:p/Sammoon391|Sammoon391]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-05-01 === * 07:24 [[phab:p/Soubag|Soubag]] was disabled by [[phab:p/Mainframe98/|Mainframe98]] === 2024-04-28 === * 06:22 [[phab:p/Diamondscoin|Diamondscoin]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-04-19 === * 03:47 [[phab:p/Wawmart2|Wawmart2]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-04-08 === * 10:21 [[phab:p/Mardetanha|Mardetanha]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2024-03-29 === * 09:03 [[phab:p/Abdollmjjedloveanan|Abdollmjjedloveanan]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-03-12 === * 09:45 [[phab:p/Samantha78462|Samantha78462]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 09:39 [[phab:p/Samantha7861654654|Samantha7861654654]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 09:11 [[phab:p/Robin|Robin]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 09:05 [[phab:p/Anglinakuki|Anglinakuki]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 00:02 [[phab:p/Johnne25|Johnne25]] was disabled by [[phab:p/bd808/|bd808]] === 2024-03-07 === * 20:07 [[phab:p/Sami785|Sami785]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-03-06 === * 07:41 [[phab:p/28q|28q]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2024-03-02 === * 22:51 [[phab:p/kitchenstrategic|kitchenstrategic]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-02-23 === * 09:03 [[phab:p/littleggghost|littleggghost]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-02-17 === * 05:10 [[phab:p/Skekeiei|Skekeiei]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-01-27 === * 03:50 [[phab:p/Andybitcoin|Andybitcoin]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2024-01-25 === * 09:10 [[phab:p/Mayo3030|Mayo3030]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-01-21 === * 05:59 [[phab:p/Hackear|Hackear]] was disabled by [[phab:p/Peachey88/|Peachey88]] * 05:58 [[phab:p/joselopez45|joselopez45]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-01-20 === * 20:02 [[phab:p/08107130655|08107130655]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2024-01-18 === * 05:33 [[phab:p/Tecnologynew|Tecnologynew]] was disabled by [[phab:p/TheresNoTime/|TheresNoTime]] === 2024-01-12 === * 22:34 [[phab:p/cchen|cchen]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2024-01-11 === * 07:34 [[phab:p/Bernita43|Bernita43]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2024-01-07 === * 10:13 [[phab:p/Irademack|Irademack]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2023-12-31 === * 16:48 [[phab:p/Vieclamdmpt|Vieclamdmpt]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2023-12-26 === * 20:34 [[phab:p/Bgu5678|Bgu5678]] was disabled by [[phab:p/Peachey88/|Peachey88]] === 2023-11-26 === * 20:09 [[phab:p/Str13tlife|Str13tlife]] was disabled by [[phab:p/JJMC89/|JJMC89]] === 2023-11-24 === * 00:49 [[phab:p/Imambuchori03|Imambuchori03]] was disabled by [[phab:p/DannyS712/|DannyS712]] === 2023-11-22 === * 15:50 [[phab:p/Naleksuh|Naleksuh]] was disabled by [[phab:p/WMFOffice/|WMFOffice]] === 2023-11-19 === * 21:10 [[phab:p/Onack16888|Onack16888]] was disabled by [[phab:p/Daimona/|Daimona]] === 2023-11-15 === * 11:57 [[phab:p/Anonymous_ehacker|Anonymous_ehacker]] was disabled by [[phab:p/hashar/|hashar]] === 2023-11-09 === * 23:16 [[phab:p/dunicorn|dunicorn]] was disabled by [[phab:p/bd808/|bd808]] === 2023-09-30 === * 17:38 [[phab:p/Wykirany|Wykirany]] was disabled by [[phab:p/RhinosF1/|RhinosF1]] === 2023-09-01 === * 16:29 [[phab:p/T200856-01|T200856-01]] was disabled by [[phab:p/bd808/|bd808]] * 16:17 [[phab:p/T200856-01|T200856-01]] was disabled by [[phab:p/bd808/|bd808]] n765kj5yqge3d4ekb17ajst8ljs3tv6 Tool:Gitlab-account-approval/Log 116 453906 2445261 2445251 2026-08-09T20:09:13Z Gitlabaccountapprovalbot 37332 yirba was rejected. 2445261 wikitext text/x-wiki <noinclude>'''Audit log of approvals''' made by [[gitlab:gitlabaccountapprovalbot|@gitlabaccountapprovalbot]]. __NOTOC__</noinclude> === 2026-08-09 === * 20:09 "yirba" was rejected (pending since 2026-05-10T20:07:07.738Z). * 06:42 "marsam2489" was rejected (pending since 2026-05-10T06:40:46.276Z). === 2026-08-08 === * 08:15 [[gitlab:taiwaniajusto|@taiwaniajusto]] was approved. === 2026-08-07 === * 18:21 [[gitlab:gturkington|@gturkington]] was approved. * 06:36 "brianbybyby" was rejected (pending since 2026-05-08T06:36:08.459Z). === 2026-07-30 === * 20:42 "horaciocolbert" was rejected (pending since 2026-04-30T20:41:40.421Z). === 2026-07-29 === * 10:03 [[gitlab:piastu|@piastu]] was approved. * 06:33 "rafiul1" was rejected (pending since 2026-04-29T06:33:02.573Z). === 2026-07-28 === * 16:51 [[gitlab:for-each-next|@for-each-next]] was approved. * 09:21 [[gitlab:gka|@gka]] was approved. === 2026-07-27 === * 08:12 [[gitlab:cambob|@cambob]] was approved. === 2026-07-26 === * 15:39 "demansanaagmailcom" was rejected (pending since 2026-04-26T15:39:03.297Z). === 2026-07-24 === * 10:54 [[gitlab:ysogo|@ysogo]] was approved. * 09:39 [[gitlab:pankaj199|@pankaj199]] was approved. === 2026-07-23 === * 12:30 [[gitlab:fermiboson|@fermiboson]] was approved. * 09:21 [[gitlab:slashme|@slashme]] was approved. === 2026-07-22 === * 15:54 [[gitlab:lmedley|@lmedley]] was approved. * 13:24 "praveen5638" was rejected (pending since 2026-04-22T13:21:23.368Z). * 12:48 [[gitlab:cyberpower678|@cyberpower678]] was approved. * 08:30 [[gitlab:nabbegat|@nabbegat]] was approved. * 08:30 [[gitlab:plyd|@plyd]] was approved. * 07:06 "ayush8620" was rejected (pending since 2026-04-22T07:03:12.476Z). === 2026-07-21 === * 15:51 [[gitlab:panieravide|@panieravide]] was approved. * 15:33 [[gitlab:luisvilla-personal|@luisvilla-personal]] was approved. * 13:48 [[gitlab:yru|@yru]] was approved. * 13:09 [[gitlab:deevad|@deevad]] was approved. * 12:54 [[gitlab:ctdo17|@ctdo17]] was approved. * 12:36 [[gitlab:jeannenoiraud|@jeannenoiraud]] was approved. * 12:21 [[gitlab:nadiantara|@nadiantara]] was approved. * 12:18 [[gitlab:wijltcher|@wijltcher]] was approved. * 10:33 [[gitlab:nivopol|@nivopol]] was approved. * 10:30 [[gitlab:johlig|@johlig]] was approved. * 10:30 [[gitlab:majicita|@majicita]] was approved. * 10:15 [[gitlab:yongjiapeng|@yongjiapeng]] was approved. * 10:09 [[gitlab:francyskus|@francyskus]] was approved. * 09:24 [[gitlab:xanonymusx|@xanonymusx]] was approved. === 2026-07-20 === * 17:15 "leonidlednev" was rejected (pending since 2026-04-20T17:13:35.108Z). * 15:27 [[gitlab:rodrigoargenton|@rodrigoargenton]] was approved. * 05:48 "draftecho" was rejected (pending since 2026-04-20T05:48:06.953Z). === 2026-07-19 === * 12:21 [[gitlab:boivie|@boivie]] was approved. === 2026-07-18 === * 16:09 [[gitlab:pharos|@pharos]] was approved. * 15:45 [[gitlab:priyankar22|@priyankar22]] was approved. * 15:30 [[gitlab:sisyph|@sisyph]] was approved. === 2026-07-13 === * 03:45 [[gitlab:dreamyshade|@dreamyshade]] was approved. === 2026-07-12 === * 09:27 [[gitlab:smk|@smk]] was approved. === 2026-07-11 === * 14:48 "bigcereal42" was rejected (pending since 2026-04-11T14:47:40.321Z). * 12:18 "pratyushsawan" was rejected (pending since 2026-04-11T12:16:41.671Z). === 2026-07-10 === * 14:42 [[gitlab:akaza24|@akaza24]] was approved. * 12:57 [[gitlab:kormisk|@kormisk]] was approved. === 2026-07-07 === * 10:36 [[gitlab:olafjanssen|@olafjanssen]] was approved. * 06:57 "elisapoly-99" was rejected (pending since 2026-04-07T06:55:52.662Z). === 2026-07-06 === * 11:57 "ma3rouf" was rejected (pending since 2026-04-06T11:56:30.978Z). === 2026-07-02 === * 15:48 [[gitlab:tekneos|@tekneos]] was approved. === 2026-07-01 === * 15:12 [[gitlab:mugurolevy|@mugurolevy]] was approved. * 14:15 [[gitlab:vadymts1|@vadymts1]] was approved. * 09:57 "mugurolevy" was rejected (pending since 2026-04-01T09:55:19.175Z). === 2026-06-30 === * 14:27 "shivangisharma" was rejected (pending since 2026-03-31T14:26:44.932Z). === 2026-06-29 === * 19:03 [[gitlab:thisismattmiller|@thisismattmiller]] was approved. === 2026-06-28 === * 14:51 "nkwenuinadine" was rejected (pending since 2026-03-29T14:48:32.735Z). * 14:03 "vaishnavikumbhar" was rejected (pending since 2026-03-29T14:01:30.604Z). * 13:03 "stepmay" was rejected (pending since 2026-03-29T13:01:59.905Z). * 06:51 "swallroth" was rejected (pending since 2026-03-29T06:49:54.838Z). === 2026-06-26 === * 09:39 [[gitlab:lakshita28|@lakshita28]] was approved. * 07:30 [[gitlab:reeti|@reeti]] was approved. * 07:30 [[gitlab:anushka10patel|@anushka10patel]] was approved. * 07:30 "samsaesque" was rejected (pending since 2026-03-27T07:29:57.279Z). * 05:51 [[gitlab:arpithhhaaa|@arpithhhaaa]] was approved. * 05:51 [[gitlab:govindlaltl|@govindlaltl]] was approved. === 2026-06-25 === * 16:09 [[gitlab:sakuraemad|@sakuraemad]] was approved. * 07:00 "kdh8219" was rejected (pending since 2026-03-26T06:58:05.415Z). === 2026-06-24 === * 11:54 [[gitlab:sanskardubeydev|@sanskardubeydev]] was approved. * 10:09 "tanmay789q" was rejected (pending since 2026-03-25T10:07:54.602Z). === 2026-06-22 === * 19:57 [[gitlab:gouvernathor|@gouvernathor]] was approved. * 16:45 [[gitlab:lucasbelo|@lucasbelo]] was approved. * 07:15 "jason2000-cpu" was rejected (pending since 2026-03-23T07:14:09.184Z). === 2026-06-21 === * 13:18 [[gitlab:egonw|@egonw]] was approved. === 2026-06-20 === * 10:21 [[gitlab:tways2017|@tways2017]] was approved. === 2026-06-19 === * 16:06 "wilsonwang2026" was rejected (pending since 2026-03-20T16:06:05.511Z). * 04:12 [[gitlab:claudio|@claudio]] was approved. === 2026-06-18 === * 14:21 "royiswariii" was rejected (pending since 2026-03-19T14:19:16.896Z). * 13:06 [[gitlab:laurabarluzzi|@laurabarluzzi]] was approved. === 2026-06-17 === * 11:24 "adinathq8x" was rejected (pending since 2026-03-18T11:22:50.098Z). * 09:45 "nathanveritas" was rejected (pending since 2026-03-18T09:43:51.645Z). === 2026-06-15 === * 22:39 [[gitlab:mohammadhijjawi|@mohammadhijjawi]] was approved. * 14:24 "enlisar" was rejected (pending since 2026-03-16T14:23:00.109Z). * 14:06 "ayaan" was rejected (pending since 2026-03-16T14:03:31.071Z). * 10:54 "kwametech" was rejected (pending since 2026-03-16T10:54:11.083Z). === 2026-06-14 === * 17:45 [[gitlab:surajseth520|@surajseth520]] was approved. * 07:24 "malahimhaseeb" was rejected (pending since 2026-03-15T07:21:57.748Z). === 2026-06-11 === * 11:48 [[gitlab:cadddr|@cadddr]] was approved. * 11:18 "wikipiggy" was rejected (pending since 2026-03-12T11:16:09.335Z). * 07:15 [[gitlab:vesihiisi|@vesihiisi]] was approved. === 2026-06-10 === * 07:03 [[gitlab:dmiranda|@dmiranda]] was approved. === 2026-06-09 === * 14:21 [[gitlab:linkgenetic|@linkgenetic]] was approved. * 14:03 [[gitlab:sjones-ctr|@sjones-ctr]] was approved. * 12:51 [[gitlab:ekrem|@ekrem]] was approved. === 2026-06-08 === * 17:48 "jmprax" was rejected (pending since 2026-03-09T17:46:38.807Z). * 16:15 [[gitlab:ahonc|@ahonc]] was approved. * 12:57 [[gitlab:rainmonger|@rainmonger]] was approved. === 2026-06-07 === * 23:03 "shadowthewuff" was rejected (pending since 2026-03-08T23:00:53.442Z). * 11:45 "wiki-pavan" was rejected (pending since 2026-03-08T11:45:11.116Z). * 02:39 [[gitlab:launchpad|@launchpad]] was approved. === 2026-06-06 === * 14:54 "unicord" was rejected (pending since 2026-03-07T14:52:04.992Z). * 12:48 "chien" was rejected (pending since 2026-03-07T12:48:11.669Z). === 2026-06-04 === * 14:33 "only-vikas" was rejected (pending since 2026-03-05T14:32:09.186Z). === 2026-06-03 === * 15:00 [[gitlab:anafibnshahibul|@anafibnshahibul]] was approved. === 2026-06-02 === * 21:21 "mgagat" was rejected (pending since 2026-03-03T21:18:37.223Z). * 13:57 "prasunaenumarthy" was rejected (pending since 2026-03-03T13:57:14.847Z). * 05:48 [[gitlab:tmoney|@tmoney]] was approved. === 2026-06-01 === * 14:57 "vikram2101" was rejected (pending since 2026-03-02T14:54:26.550Z). * 12:03 "watshell" was rejected (pending since 2026-03-02T12:03:09.329Z). === 2026-05-29 === * 12:48 "mounikapotladurthi" was rejected (pending since 2026-02-27T12:45:38.609Z). === 2026-05-27 === * 20:00 "vinitha" was rejected (pending since 2026-02-25T19:58:43.524Z). * 16:30 "codeurluce" was rejected (pending since 2026-02-25T16:28:53.973Z). * 14:33 [[gitlab:thilio|@thilio]] was approved. === 2026-05-26 === * 12:09 "charisad" was rejected (pending since 2026-02-24T12:07:21.881Z). === 2026-05-25 === * 22:54 "ddshelto" was rejected (pending since 2026-02-23T22:52:44.427Z). * 19:51 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z). * 19:48 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z). === 2026-05-24 === * 18:45 "jiyagupta-cs" was rejected (pending since 2026-02-22T18:43:33.176Z). === 2026-05-23 === * 13:09 [[gitlab:gauthammohanraj|@gauthammohanraj]] was approved. * 04:21 [[gitlab:staraction|@staraction]] was approved. === 2026-05-22 === * 19:03 "i-horich" was rejected (pending since 2026-02-20T19:00:43.519Z). * 01:48 "50323233" was rejected (pending since 2026-02-20T01:48:05.555Z). === 2026-05-21 === * 18:51 "kartikeyg0104" was rejected (pending since 2026-02-19T18:48:39.707Z). * 16:27 [[gitlab:renovatebot|@renovatebot]] was approved. * 16:06 [[gitlab:gkm563|@gkm563]] was approved. === 2026-05-20 === * 01:21 "beedellrokejulianlockhart" was rejected (pending since 2026-02-18T01:19:13.284Z). === 2026-05-18 === * 23:18 "wladek92" was rejected (pending since 2026-02-16T23:16:22.939Z). * 16:36 [[gitlab:effeietsanders|@effeietsanders]] was approved. === 2026-05-14 === * 21:00 [[gitlab:nehemienathan|@nehemienathan]] was approved. === 2026-05-13 === * 10:51 "ssssaaaa" was rejected (pending since 2026-02-11T10:50:36.975Z). === 2026-05-12 === * 18:06 [[gitlab:psubhashish|@psubhashish]] was approved. * 08:12 "khan" was rejected (pending since 2026-02-10T08:11:48.776Z). * 04:27 "galaxysh" was rejected (pending since 2026-02-10T04:24:59.440Z). === 2026-05-11 === * 12:18 "peterxy12" was rejected (pending since 2026-02-09T12:18:01.982Z). === 2026-05-10 === * 11:09 "yalihupokn" was rejected (pending since 2026-02-08T11:06:51.336Z). * 05:12 "wobadha" was rejected (pending since 2026-02-08T05:11:00.569Z). === 2026-05-09 === * 13:45 "bwiki" was rejected (pending since 2026-02-07T13:43:38.177Z). === 2026-05-08 === * 09:24 [[gitlab:cwilliams|@cwilliams]] was approved. === 2026-05-07 === * 14:15 "rehankhan78" was rejected (pending since 2026-02-05T14:13:37.754Z). === 2026-05-06 === * 11:24 "ari" was rejected (pending since 2026-02-04T11:24:11.760Z). * 08:09 [[gitlab:neriah|@neriah]] was approved. * 06:27 [[gitlab:status401|@status401]] was approved. === 2026-05-03 === * 09:54 [[gitlab:anilk|@anilk]] was approved. === 2026-05-02 === * 17:54 [[gitlab:sweil|@sweil]] was approved. * 17:00 [[gitlab:aoppo|@aoppo]] was approved. === 2026-05-01 === * 21:18 [[gitlab:dawalda|@dawalda]] was approved. === 2026-04-30 === * 21:42 "merohibine" was rejected (pending since 2026-01-29T21:40:00.756Z). * 20:54 [[gitlab:tfmorris|@tfmorris]] was approved. * 17:33 [[gitlab:uyen|@uyen]] was approved. * 07:39 [[gitlab:mahveotm|@mahveotm]] was approved. * 06:36 [[gitlab:leo321|@leo321]] was approved. === 2026-04-29 === * 02:27 [[gitlab:dw31415|@dw31415]] was approved. === 2026-04-28 === * 23:09 [[gitlab:dtorsani|@dtorsani]] was approved. === 2026-04-27 === * 23:42 [[gitlab:quinlan|@quinlan]] was approved. * 05:00 [[gitlab:matthewyeager|@matthewyeager]] was approved. === 2026-04-26 === * 17:36 "kuba-hajnej" was rejected (pending since 2026-01-25T17:33:32.467Z). * 13:03 "jklamo" was rejected (pending since 2026-01-25T13:02:22.936Z). === 2026-04-25 === * 20:24 [[gitlab:maldaxura|@maldaxura]] was approved. * 14:33 [[gitlab:sirtobi|@sirtobi]] was approved. * 04:18 "ice5678" was rejected (pending since 2026-01-24T04:15:30.008Z). === 2026-04-24 === * 22:06 [[gitlab:arcstur|@arcstur]] was approved. === 2026-04-22 === * 23:06 "dtorsani" was rejected (pending since 2026-01-21T23:03:25.843Z). * 22:18 [[gitlab:egezort|@egezort]] was approved. * 16:45 "nexpectarpit" was rejected (pending since 2026-01-21T16:43:21.045Z). === 2026-04-20 === * 19:15 "fitch" was rejected (pending since 2026-01-19T19:12:35.644Z). === 2026-04-19 === * 02:54 [[gitlab:neoact|@neoact]] was approved. === 2026-04-18 === * 07:06 [[gitlab:kockaadmiralac|@kockaadmiralac]] was approved. === 2026-04-17 === * 13:42 "liselot" was rejected (pending since 2026-01-16T13:39:41.909Z). === 2026-04-15 === * 17:03 "lahari" was rejected (pending since 2026-01-14T17:02:06.275Z). === 2026-04-14 === * 13:00 "surajseth520" was rejected (pending since 2026-01-13T12:59:45.906Z). * 04:51 [[gitlab:canley|@canley]] was approved. * 01:03 "bshizzle" was rejected (pending since 2026-01-13T01:00:48.120Z). === 2026-04-13 === * 15:30 [[gitlab:passimacopoulos|@passimacopoulos]] was approved. === 2026-04-11 === * 12:30 "krithash" was rejected (pending since 2026-01-10T12:27:24.731Z). === 2026-04-10 === * 15:30 "raunak1709" was rejected (pending since 2026-01-09T15:29:10.901Z). === 2026-04-07 === * 17:03 [[gitlab:supnabla|@supnabla]] was approved. === 2026-04-06 === * 20:00 [[gitlab:laerdon|@laerdon]] was approved. * 19:21 [[gitlab:ljq3|@ljq3]] was approved. === 2026-04-04 === * 11:06 "mixcc" was rejected (pending since 2026-01-03T11:03:33.922Z). === 2026-04-02 === * 05:30 [[gitlab:mbh1|@mbh1]] was approved. === 2026-04-01 === * 18:21 "yuvrajpatil17" was rejected (pending since 2025-12-31T18:20:27.991Z). * 12:12 [[gitlab:amorii0|@amorii0]] was approved. === 2026-03-31 === * 11:00 "krrishsehgal" was rejected (pending since 2025-12-30T11:00:16.384Z). === 2026-03-30 === * 15:36 [[gitlab:atsuko|@atsuko]] was approved. === 2026-03-29 === * 11:36 [[gitlab:giftcup|@giftcup]] was approved. === 2026-03-28 === * 14:51 [[gitlab:janeeva1|@janeeva1]] was approved. === 2026-03-26 === * 13:36 [[gitlab:saiphani02|@saiphani02]] was approved. * 11:48 [[gitlab:valerioboz-wmch|@valerioboz-wmch]] was approved. === 2026-03-25 === * 09:45 "quansi" was rejected (pending since 2025-12-24T09:42:13.451Z). * 02:18 [[gitlab:viztor|@viztor]] was approved. === 2026-03-24 === * 23:18 [[gitlab:maryyann|@maryyann]] was approved. * 23:01 [[gitlab:codenamenoreste|@codenamenoreste]] was approved. * 13:36 [[gitlab:marc-maillard-wmse|@marc-maillard-wmse]] was approved. * 07:39 "fred2675" was rejected (pending since 2025-12-23T07:39:11.380Z). === 2026-03-23 === * 14:51 [[gitlab:komla|@komla]] was approved. * 05:51 "lunachuck43" was rejected (pending since 2025-12-22T05:50:17.862Z). * 04:06 "reza110011" was rejected (pending since 2025-12-22T04:05:25.117Z). === 2026-03-20 === * 21:54 "mertgor" was rejected (pending since 2025-12-19T21:51:51.419Z). * 20:57 "autanmahmah" was rejected (pending since 2025-12-19T20:54:51.678Z). * 09:57 [[gitlab:nethahussain|@nethahussain]] was approved. * 09:27 [[gitlab:piewriter|@piewriter]] was approved. * 08:15 [[gitlab:dondersmooi|@dondersmooi]] was approved. === 2026-03-19 === * 21:03 "sayvhior" was rejected (pending since 2025-12-18T21:02:31.699Z). === 2026-03-18 === * 20:15 [[gitlab:martinmystere|@martinmystere]] was approved. === 2026-03-17 === * 02:51 "louperivois" was rejected (pending since 2025-12-16T02:50:48.197Z). === 2026-03-16 === * 12:54 "mokayaj857" was rejected (pending since 2025-12-15T12:53:39.015Z). * 06:18 "roamer15" was rejected (pending since 2025-12-15T06:16:38.042Z). === 2026-03-14 === * 11:12 "umaramuhammad" was rejected (pending since 2025-12-13T11:10:44.004Z). * 09:33 "akuma19" was rejected (pending since 2025-12-13T09:31:39.044Z). * 07:06 [[gitlab:syunsyunminmin|@syunsyunminmin]] was approved. === 2026-03-12 === * 20:24 [[gitlab:11wb|@11wb]] was approved. * 09:54 [[gitlab:bcxfu75k|@bcxfu75k]] was approved. === 2026-03-10 === * 09:12 [[gitlab:viktoriahillerudwmse|@viktoriahillerudwmse]] was approved. === 2026-03-06 === * 08:09 "vazhayilnewone" was rejected (pending since 2025-12-05T08:07:02.184Z). === 2026-03-04 === * 20:54 [[gitlab:elphie|@elphie]] was approved. * 11:39 "ronaldahmed" was rejected (pending since 2025-12-03T11:37:47.492Z). * 02:12 "ltslw" was rejected (pending since 2025-12-03T02:11:52.040Z). === 2026-03-02 === * 19:21 "dlopez350" was rejected (pending since 2025-12-01T19:20:38.918Z). * 18:15 [[gitlab:lsandergreen|@lsandergreen]] was approved. === 2026-03-01 === * 10:51 [[gitlab:clintacc|@clintacc]] was approved. === 2026-02-28 === * 09:24 "cardboardlamp" was rejected (pending since 2025-11-29T09:22:03.947Z). * 08:18 "wiki-pavan" was rejected (pending since 2025-11-29T08:16:24.184Z). === 2026-02-27 === * 20:45 "thisisrick25" was rejected (pending since 2025-11-28T20:42:24.454Z). === 2026-02-26 === * 13:57 "chuiimuiiofc" was rejected (pending since 2025-11-27T13:57:02.794Z). * 13:54 "steffpro" was rejected (pending since 2025-11-27T13:52:10.859Z). === 2026-02-25 === * 21:24 "abubakarhabibudayyabu" was rejected (pending since 2025-11-26T21:22:37.776Z). === 2026-02-24 === * 05:00 "playboi" was rejected (pending since 2025-11-25T05:00:30.762Z). === 2026-02-23 === * 14:00 "alph65" was rejected (pending since 2025-11-24T13:59:00.797Z). * 12:33 [[gitlab:robertsky|@robertsky]] was approved. === 2026-02-22 === * 00:30 "hp8p" was rejected (pending since 2025-11-23T00:29:24.741Z). === 2026-02-19 === * 16:45 "clayjar" was rejected (pending since 2025-11-20T16:44:48.380Z). === 2026-02-18 === * 22:18 "nexus" was rejected (pending since 2025-11-19T22:16:48.818Z). * 12:00 "bernsteinnn" was rejected (pending since 2025-11-19T11:59:04.427Z). === 2026-02-17 === * 11:36 "jason2000-cpu" was rejected (pending since 2025-11-18T11:34:00.314Z). === 2026-02-16 === * 14:54 "smaurya" was rejected (pending since 2025-11-17T14:52:06.906Z). === 2026-02-15 === * 16:51 "kra-79" was rejected (pending since 2025-11-16T16:50:41.375Z). === 2026-02-14 === * 15:15 [[gitlab:mess|@mess]] was approved. === 2026-02-13 === * 13:57 "sopalsuemae957" was rejected (pending since 2025-11-14T13:55:16.921Z). * 13:30 [[gitlab:wyslijp16-toolforge|@wyslijp16-toolforge]] was approved. === 2026-02-12 === * 16:30 "kristinagligoric" was rejected (pending since 2025-11-13T16:29:21.646Z). * 03:33 [[gitlab:anyehansen|@anyehansen]] was approved. * 02:21 [[gitlab:thejoyfultentmaker|@thejoyfultentmaker]] was approved. === 2026-02-10 === * 13:18 [[gitlab:db111|@db111]] was approved. === 2026-02-09 === * 19:06 "squirrel289" was rejected (pending since 2025-11-10T19:04:27.831Z). === 2026-02-06 === * 20:54 [[gitlab:gillux|@gillux]] was approved. * 09:09 [[gitlab:lih|@lih]] was approved. === 2026-01-31 === * 16:21 [[gitlab:taxonbot1|@taxonbot1]] was approved. === 2026-01-28 === * 14:30 [[gitlab:ademola|@ademola]] was approved. * 10:51 "watshell" was rejected (pending since 2025-10-29T10:51:01.521Z). === 2026-01-26 === * 23:06 "tavaresgmg" was rejected (pending since 2025-10-27T23:04:42.140Z). === 2026-01-25 === * 06:03 "cata" was rejected (pending since 2025-10-26T06:01:26.155Z). === 2026-01-24 === * 21:15 [[gitlab:wiegels|@wiegels]] was approved. * 06:30 [[gitlab:blaquans|@blaquans]] was approved. === 2026-01-23 === * 16:27 [[gitlab:lerickson|@lerickson]] was approved. * 10:15 "fran0035g" was rejected (pending since 2025-10-24T10:12:17.732Z). === 2026-01-22 === * 21:00 "hacksyn" was rejected (pending since 2025-10-23T20:59:15.982Z). === 2026-01-21 === * 17:30 [[gitlab:otcenas11|@otcenas11]] was approved. === 2026-01-19 === * 21:48 [[gitlab:amdrel|@amdrel]] was approved. * 04:36 "rayalexa" was rejected (pending since 2025-10-20T04:35:02.094Z). === 2026-01-18 === * 15:45 "somya" was rejected (pending since 2025-10-19T15:43:43.701Z). * 06:54 "sergg001" was rejected (pending since 2025-10-19T06:54:12.296Z). === 2026-01-16 === * 11:57 "zeejohsy" was rejected (pending since 2025-10-17T11:56:22.372Z). * 04:45 "rocky25" was rejected (pending since 2025-10-17T04:43:33.180Z). === 2026-01-15 === * 16:39 "tiisu" was rejected (pending since 2025-10-16T16:37:18.438Z). * 12:00 "noahalorwu" was rejected (pending since 2025-10-16T11:58:26.133Z). * 10:39 "prjayaiuedu" was rejected (pending since 2025-10-16T10:37:16.947Z). === 2026-01-13 === * 17:21 [[gitlab:lwilson-ctr|@lwilson-ctr]] was approved. === 2026-01-12 === * 17:03 "stagietechs" was rejected (pending since 2025-10-13T17:02:25.281Z). === 2026-01-10 === * 19:06 "keerthisr" was rejected (pending since 2025-10-11T19:05:01.758Z). === 2026-01-09 === * 20:36 "lightb" was rejected (pending since 2025-10-10T20:34:20.264Z). === 2026-01-08 === * 19:42 [[gitlab:tbodt|@tbodt]] was approved. * 13:57 [[gitlab:martynranyard|@martynranyard]] was approved. === 2026-01-07 === * 17:48 [[gitlab:santanuwiki25|@santanuwiki25]] was approved. * 14:27 "dipanshu" was rejected (pending since 2025-10-08T14:26:10.794Z). * 12:30 "adeolaadesina" was rejected (pending since 2025-10-08T12:29:49.592Z). * 09:21 "tony-kamande" was rejected (pending since 2025-10-08T09:20:28.421Z). * 06:18 "hninwuttyi" was rejected (pending since 2025-10-08T06:17:28.006Z). * 05:09 "andume" was rejected (pending since 2025-10-08T05:07:18.582Z). * 02:00 "mosope" was rejected (pending since 2025-10-08T01:59:54.800Z). * 01:15 [[gitlab:tungstalite|@tungstalite]] was approved. === 2026-01-06 === * 18:24 "leerensucher" was rejected (pending since 2025-10-07T18:21:41.253Z). * 14:54 "leonidlednev" was rejected (pending since 2025-10-07T14:53:07.273Z). * 12:57 "alexandre-tingaud" was rejected (pending since 2025-10-07T12:54:27.206Z). === 2026-01-04 === * 21:33 [[gitlab:matr1x-101|@matr1x-101]] was approved. * 15:18 "makjr" was rejected (pending since 2025-10-05T15:16:31.558Z). * 14:09 "dakshq" was rejected (pending since 2025-10-05T14:08:40.608Z). === 2026-01-03 === * 20:42 [[gitlab:apehitkey|@apehitkey]] was approved. * 18:00 [[gitlab:jeremyb|@jeremyb]] was approved. * 14:09 [[gitlab:twelephant|@twelephant]] was approved. === 2026-01-01 === * 11:30 "shellstanislav" was rejected (pending since 2025-10-02T11:29:10.150Z). === 2025-12-30 === * 19:51 "camilojdiaz" was rejected (pending since 2025-09-30T19:49:24.913Z). === 2025-12-29 === * 16:03 "zied" was rejected (pending since 2025-09-29T16:01:30.415Z). * 08:18 "rahulsidpradhan" was rejected (pending since 2025-09-29T08:17:02.849Z). === 2025-12-26 === * 09:48 "thembo42" was rejected (pending since 2025-09-26T09:45:15.033Z). === 2025-12-25 === * 14:03 "196936074751" was rejected (pending since 2025-09-25T14:02:31.367Z). === 2025-12-23 === * 16:21 "ngarnsworthy" was rejected (pending since 2025-09-23T16:20:41.211Z). === 2025-12-22 === * 12:39 "aza555" was rejected (pending since 2025-09-22T12:38:02.622Z). === 2025-12-20 === * 23:45 "saph" was rejected (pending since 2025-09-20T23:45:01.222Z). === 2025-12-19 === * 10:15 "vladdymoses" was rejected (pending since 2025-09-19T10:15:00.999Z). * 07:15 "dirtylittlepoobah" was rejected (pending since 2025-09-19T07:13:55.537Z). === 2025-12-18 === * 16:24 [[gitlab:guyfawcus|@guyfawcus]] was approved. === 2025-12-17 === * 21:39 [[gitlab:holdyourhorses|@holdyourhorses]] was approved. * 18:30 "prudencia" was rejected (pending since 2025-09-17T18:27:18.860Z). * 02:24 "lottie" was rejected (pending since 2025-09-17T02:21:21.744Z). === 2025-12-16 === * 09:39 [[gitlab:melcatherine|@melcatherine]] was approved. * 08:54 [[gitlab:leila237|@leila237]] was approved. === 2025-12-15 === * 18:27 [[gitlab:royalsailor|@royalsailor]] was approved. * 09:39 [[gitlab:olaf8940|@olaf8940]] was approved. * 09:39 "brianbybyby" was rejected (pending since 2025-09-15T09:37:45.430Z). === 2025-12-14 === * 20:21 [[gitlab:essa237|@essa237]] was approved. * 16:42 [[gitlab:bovimacoco|@bovimacoco]] was approved. === 2025-12-13 === * 21:54 "mmns21" was rejected (pending since 2025-09-13T21:52:24.017Z). * 20:33 "bugcrawler" was rejected (pending since 2025-09-13T20:31:09.211Z). === 2025-12-12 === * 14:39 "ruvchoudhary" was rejected (pending since 2025-09-12T14:36:16.167Z). * 06:54 "rezadress" was rejected (pending since 2025-09-12T06:52:21.749Z). === 2025-12-10 === * 17:30 [[gitlab:itsmoon|@itsmoon]] was approved. === 2025-12-09 === * 15:42 [[gitlab:mercy-o|@mercy-o]] was approved. === 2025-12-06 === * 16:45 "jacquesradjabu" was rejected (pending since 2025-09-06T16:45:17.969Z). * 11:27 [[gitlab:ikhitron|@ikhitron]] was approved. === 2025-12-01 === * 08:12 "halconmilenario21" was rejected (pending since 2025-09-01T08:12:10.262Z). === 2025-11-30 === * 21:06 [[gitlab:habs|@habs]] was approved. === 2025-11-29 === * 16:36 "bovimacoco" was rejected (pending since 2025-08-30T16:34:39.712Z). * 00:45 [[gitlab:jjpmaster|@jjpmaster]] was approved. === 2025-11-24 === * 10:30 "alph65" was rejected (pending since 2025-08-25T10:28:40.957Z). * 02:24 [[gitlab:yaron|@yaron]] was approved. === 2025-11-20 === * 16:06 "clayjar" was rejected (pending since 2025-08-21T16:04:54.450Z). === 2025-11-17 === * 21:09 [[gitlab:ankita97531|@ankita97531]] was approved. === 2025-11-16 === * 14:15 "commanderkefir" was rejected (pending since 2025-08-17T14:13:14.791Z). * 08:21 "rehankhan78" was rejected (pending since 2025-08-17T08:19:44.896Z). === 2025-11-15 === * 14:36 "cyberscribe" was rejected (pending since 2025-08-16T14:34:27.230Z). === 2025-11-13 === * 04:21 "waddie96" was rejected (pending since 2025-08-14T04:19:27.461Z). === 2025-11-11 === * 06:42 [[gitlab:seanhoyland|@seanhoyland]] was approved. === 2025-11-10 === * 00:06 [[gitlab:jaredblumer|@jaredblumer]] was approved. === 2025-11-09 === * 22:36 "heinxiety" was rejected (pending since 2025-08-10T22:33:12.041Z). === 2025-11-07 === * 22:00 [[gitlab:forzagreen|@forzagreen]] was approved. === 2025-11-06 === * 16:57 [[gitlab:rsilvola|@rsilvola]] was approved. === 2025-11-04 === * 21:24 [[gitlab:devdoingdev|@devdoingdev]] was approved. === 2025-11-03 === * 17:48 "joewaleed98" was rejected (pending since 2025-08-04T17:46:12.191Z). === 2025-11-01 === * 18:00 "eliasempresas" was rejected (pending since 2025-08-02T17:58:04.412Z). === 2025-10-31 === * 18:51 [[gitlab:chaoticenby|@chaoticenby]] was approved. * 04:33 "3ch310n" was rejected (pending since 2025-08-01T04:32:21.982Z). === 2025-10-30 === * 10:03 [[gitlab:tausheefhassan|@tausheefhassan]] was approved. === 2025-10-29 === * 14:54 "theap" was rejected (pending since 2025-07-30T14:52:12.066Z). === 2025-10-28 === * 06:06 [[gitlab:tanbiruzzaman|@tanbiruzzaman]] was approved. === 2025-10-27 === * 07:51 [[gitlab:jmoore111|@jmoore111]] was approved. === 2025-10-25 === * 21:09 [[gitlab:valor|@valor]] was approved. * 21:03 [[gitlab:booksmurf|@booksmurf]] was approved. * 02:48 "mystyc1" was rejected (pending since 2025-07-26T02:46:19.373Z). === 2025-10-24 === * 05:12 "aadarshmahesh" was rejected (pending since 2025-07-25T05:09:38.264Z). === 2025-10-22 === * 20:54 [[gitlab:janewanga|@janewanga]] was approved. * 17:27 "abeljeevan" was rejected (pending since 2025-07-23T17:26:46.884Z). * 16:12 "shrimpnaur" was rejected (pending since 2025-07-23T16:10:37.864Z). === 2025-10-21 === * 18:51 "jrmuizel" was rejected (pending since 2025-07-22T18:50:07.315Z). * 09:33 [[gitlab:dpogorzelski|@dpogorzelski]] was approved. === 2025-10-17 === * 13:21 [[gitlab:blegodwin|@blegodwin]] was approved. === 2025-10-16 === * 14:51 [[gitlab:bahago|@bahago]] was approved. * 14:12 "harikrishna0005" was rejected (pending since 2025-07-17T14:10:48.385Z). * 14:09 "gauthammohanraj" was rejected (pending since 2025-07-17T14:08:47.643Z). === 2025-10-15 === * 13:48 [[gitlab:adwivedii|@adwivedii]] was approved. * 13:18 [[gitlab:kimbrenekakande|@kimbrenekakande]] was approved. * 13:03 "childmnajennifer" was rejected (pending since 2025-07-16T13:01:50.236Z). * 05:06 "vssb4214" was rejected (pending since 2025-07-16T05:05:33.985Z). === 2025-10-14 === * 19:39 [[gitlab:afanyulionel|@afanyulionel]] was approved. * 15:33 [[gitlab:sadrettin|@sadrettin]] was approved. * 14:18 [[gitlab:tmwyk|@tmwyk]] was approved. * 08:42 "yasu0796" was rejected (pending since 2025-07-15T08:41:26.453Z). === 2025-10-13 === * 16:09 [[gitlab:atlas0007|@atlas0007]] was approved. === 2025-10-11 === * 17:42 [[gitlab:techwizzie|@techwizzie]] was approved. === 2025-10-10 === * 19:03 [[gitlab:miiswom|@miiswom]] was approved. * 16:06 [[gitlab:ninatakang|@ninatakang]] was approved. === 2025-10-09 === * 15:42 [[gitlab:jaykaneki|@jaykaneki]] was approved. * 14:21 [[gitlab:lebogang|@lebogang]] was approved. * 14:15 [[gitlab:kimondorose|@kimondorose]] was approved. * 13:48 [[gitlab:joyakinyi|@joyakinyi]] was approved. * 13:48 [[gitlab:dikshyashahi|@dikshyashahi]] was approved. * 13:45 [[gitlab:obediobadiah|@obediobadiah]] was approved. * 13:45 [[gitlab:system625|@system625]] was approved. * 13:45 [[gitlab:rolalove|@rolalove]] was approved. * 13:39 [[gitlab:olatundeawo|@olatundeawo]] was approved. * 13:36 [[gitlab:danielchristlight|@danielchristlight]] was approved. * 13:36 [[gitlab:dipanshu1223|@dipanshu1223]] was approved. * 13:36 [[gitlab:aradhya|@aradhya]] was approved. * 09:57 "bognd" was rejected (pending since 2025-07-10T09:55:48.661Z). === 2025-10-08 === * 23:36 [[gitlab:sopzy|@sopzy]] was approved. * 23:03 [[gitlab:oluwatumininu|@oluwatumininu]] was approved. * 19:39 [[gitlab:levon003|@levon003]] was approved. * 15:24 [[gitlab:ritika-bhambri11|@ritika-bhambri11]] was approved. * 13:45 [[gitlab:anbanguyen|@anbanguyen]] was approved. * 13:36 [[gitlab:chumzine|@chumzine]] was approved. * 13:27 [[gitlab:shr0x-ya|@shr0x-ya]] was approved. * 12:45 [[gitlab:nurahwakili|@nurahwakili]] was approved. * 03:42 "nazhiba" was rejected (pending since 2025-07-09T03:40:12.625Z). * 02:12 "mafennel" was rejected (pending since 2025-07-09T02:11:40.598Z). === 2025-10-07 === * 22:54 [[gitlab:olusegunfaj|@olusegunfaj]] was approved. * 21:30 [[gitlab:rona|@rona]] was approved. * 21:09 [[gitlab:sandijigs|@sandijigs]] was approved. * 13:36 "xisbajao" was rejected (pending since 2025-07-08T13:33:35.018Z). * 01:36 "areczek94" was rejected (pending since 2025-07-08T01:35:40.633Z). === 2025-10-06 === * 19:21 "wmcarter2017" was rejected (pending since 2025-07-07T19:21:12.899Z). === 2025-10-05 === * 14:15 "meetmendapara" was rejected (pending since 2025-07-06T14:14:16.726Z). === 2025-10-04 === * 20:51 "nftbaee" was rejected (pending since 2025-07-05T20:50:57.688Z). === 2025-10-03 === * 06:12 [[gitlab:javiermonton|@javiermonton]] was approved. === 2025-10-02 === * 20:15 "talaqalotaibipmp" was rejected (pending since 2025-07-03T20:13:05.164Z). === 2025-10-01 === * 10:54 "bjensen" was rejected (pending since 2025-07-02T10:53:46.574Z). * 02:45 "kowal1984" was rejected (pending since 2025-07-02T02:44:56.946Z). === 2025-09-30 === * 21:21 [[gitlab:kavaljeetsingh|@kavaljeetsingh]] was approved. * 00:24 "adium" was rejected (pending since 2025-07-01T00:23:43.807Z). === 2025-09-28 === * 08:54 [[gitlab:pexerik|@pexerik]] was approved. === 2025-09-27 === * 13:57 [[gitlab:rubahhitamvukova|@rubahhitamvukova]] was approved. === 2025-09-26 === * 16:57 "algorithmic" was rejected (pending since 2025-06-27T16:56:17.480Z). * 13:54 [[gitlab:shadabgdg|@shadabgdg]] was approved. * 13:12 [[gitlab:spushpit|@spushpit]] was approved. === 2025-09-20 === * 14:06 "bwiki" was rejected (pending since 2025-06-21T13:59:14.749Z). === 2025-09-16 === * 05:39 [[gitlab:deepchirp|@deepchirp]] was approved. === 2025-09-15 === * 22:00 [[gitlab:noisk8|@noisk8]] was approved. * 11:03 "ahonc" was rejected (pending since 2025-06-16T11:00:54.843Z). === 2025-09-13 === * 18:24 "a-ssh22" was rejected (pending since 2025-06-14T18:23:33.937Z). * 12:36 [[gitlab:rajashreetalukdar|@rajashreetalukdar]] was approved. * 00:45 [[gitlab:sumitsurai|@sumitsurai]] was approved. === 2025-09-12 === * 17:12 [[gitlab:suyash23|@suyash23]] was approved. * 00:46 "remotetravel" was rejected (pending since 2025-06-13T00:44:08.171Z). === 2025-09-10 === * 21:09 "jancborchardt" was rejected (pending since 2025-06-11T21:06:30.759Z). === 2025-09-09 === * 17:03 [[gitlab:vwf|@vwf]] was approved. * 06:36 [[gitlab:cactusisme|@cactusisme]] was approved. === 2025-09-08 === * 18:09 "birushandegeya" was rejected (pending since 2025-06-09T18:08:00.087Z). * 16:27 "ngarnsworthy" was rejected (pending since 2025-06-09T16:24:37.213Z). * 12:33 "zolgoyo" was rejected (pending since 2025-06-09T12:31:34.199Z). === 2025-09-06 === * 23:09 [[gitlab:jaishsingh913|@jaishsingh913]] was approved. === 2025-09-05 === * 21:45 [[gitlab:sakshi2|@sakshi2]] was approved. * 20:42 "abdukhaliq1" was rejected (pending since 2025-06-06T20:40:42.023Z). * 14:27 "beubsamy" was rejected (pending since 2025-06-06T14:27:06.781Z). === 2025-09-04 === * 23:27 "sdhehua" was rejected (pending since 2025-06-05T23:24:45.777Z). * 19:00 [[gitlab:perry|@perry]] was approved. * 11:24 "saintwolf" was rejected (pending since 2025-06-05T11:21:20.176Z). === 2025-09-02 === * 05:48 [[gitlab:aliu|@aliu]] was approved. === 2025-08-29 === * 13:30 "kksurendran066" was rejected (pending since 2025-05-30T13:27:48.755Z). === 2025-08-28 === * 22:18 "tauraamuix" was rejected (pending since 2025-05-29T22:16:08.228Z). === 2025-08-26 === * 19:03 [[gitlab:dikkulah|@dikkulah]] was approved. === 2025-08-22 === * 23:51 [[gitlab:khoroshun_mike|@khoroshun_mike]] was approved. === 2025-08-21 === * 07:39 [[gitlab:yuka|@yuka]] was approved. === 2025-08-19 === * 07:48 [[gitlab:zhaofjx|@zhaofjx]] was approved. === 2025-08-17 === * 14:27 "madhan13k" was rejected (pending since 2025-05-18T14:26:08.973Z). === 2025-08-15 === * 10:15 "mohammed_abukhadra" was rejected (pending since 2025-05-16T10:14:48.403Z). === 2025-08-11 === * 11:48 "hmmyesbro" was rejected (pending since 2025-05-12T11:45:24.350Z). === 2025-08-10 === * 13:15 [[gitlab:dactyl|@dactyl]] was approved. === 2025-08-09 === * 04:39 "xxxx100000" was rejected (pending since 2025-05-10T04:37:44.949Z). === 2025-08-08 === * 14:33 [[gitlab:josefanthony|@josefanthony]] was approved. === 2025-08-07 === * 23:42 [[gitlab:robins7|@robins7]] was approved. * 21:42 [[gitlab:pols12|@pols12]] was approved. * 17:15 "sbronson" was rejected (pending since 2025-05-08T17:15:08.834Z). * 14:57 [[gitlab:alvindulle|@alvindulle]] was approved. * 14:45 [[gitlab:xentos|@xentos]] was approved. * 06:27 "jamesboste" was rejected (pending since 2025-05-08T06:25:14.793Z). * 03:57 "ysun" was rejected (pending since 2025-05-08T03:55:07.348Z). === 2025-08-06 === * 21:51 "pols12" was rejected (pending since 2025-05-07T21:49:13.598Z). * 01:51 "okeamah" was rejected (pending since 2025-05-07T01:48:50.114Z). === 2025-08-05 === * 09:15 "mobashir-2013" was rejected (pending since 2025-05-06T09:14:24.069Z). === 2025-08-01 === * 08:00 "douginamug" was rejected (pending since 2025-05-02T07:57:38.317Z). === 2025-07-31 === * 02:30 [[gitlab:ads|@ads]] was approved. === 2025-07-27 === * 13:15 "mrico2703" was rejected (pending since 2025-04-27T13:13:12.346Z). * 10:17 [[gitlab:josephfrancis12|@josephfrancis12]] was approved. * 10:17 [[gitlab:fuzzew|@fuzzew]] was approved. * 05:57 [[gitlab:biscuitbobby|@biscuitbobby]] was approved. * 05:48 [[gitlab:ecoholic|@ecoholic]] was approved. === 2025-07-26 === * 11:48 [[gitlab:chimnayyyy|@chimnayyyy]] was approved. * 11:48 [[gitlab:alwinalbert|@alwinalbert]] was approved. * 11:48 [[gitlab:hridyakk|@hridyakk]] was approved. * 11:45 [[gitlab:gaurigupta21|@gaurigupta21]] was approved. * 11:45 [[gitlab:binetaa|@binetaa]] was approved. * 10:21 [[gitlab:jyothikat22|@jyothikat22]] was approved. * 10:21 [[gitlab:zobotrombie|@zobotrombie]] was approved. * 10:21 [[gitlab:flykrth|@flykrth]] was approved. * 10:21 [[gitlab:mehrinshamim|@mehrinshamim]] was approved. * 10:21 [[gitlab:aadhi13|@aadhi13]] was approved. * 10:21 [[gitlab:malavikam05|@malavikam05]] was approved. * 10:18 [[gitlab:nf609|@nf609]] was approved. * 05:48 [[gitlab:nazalnihad|@nazalnihad]] was approved. * 05:48 [[gitlab:naveen28204280|@naveen28204280]] was approved. === 2025-07-25 === * 09:49 [[gitlab:kasyap9|@kasyap9]] was approved. * 09:30 [[gitlab:swayamagrahari|@swayamagrahari]] was approved. === 2025-07-24 === * 19:36 [[gitlab:madutgn|@madutgn]] was approved. === 2025-07-23 === * 20:09 [[gitlab:somerandomdeveloper|@somerandomdeveloper]] was approved. === 2025-07-22 === * 00:15 [[gitlab:iagoqnsi|@iagoqnsi]] was approved. === 2025-07-21 === * 17:30 [[gitlab:asadiqui|@asadiqui]] was approved. * 16:39 [[gitlab:tryvix1509|@tryvix1509]] was approved. * 04:27 [[gitlab:damian|@damian]] was approved. === 2025-07-20 === * 09:42 "mike-khoroshun" was rejected (pending since 2025-04-20T09:42:22.732Z). === 2025-07-17 === * 17:57 [[gitlab:haroldkrabs|@haroldkrabs]] was approved. * 13:45 [[gitlab:envlh|@envlh]] was approved. === 2025-07-14 === * 10:24 [[gitlab:missguru|@missguru]] was approved. * 00:57 "clarfonthey" was rejected (pending since 2025-04-14T00:56:32.626Z). === 2025-07-13 === * 01:01 [[gitlab:l235|@l235]] was approved. === 2025-07-11 === * 03:06 "rodavlas" was rejected (pending since 2025-04-11T03:05:45.590Z). === 2025-07-06 === * 00:09 "lakasa" was rejected (pending since 2025-04-06T00:06:28.469Z). === 2025-07-05 === * 21:54 "ctrlzvi" was rejected (pending since 2025-04-05T21:54:12.542Z). * 14:30 "aminualiyu" was rejected (pending since 2025-04-05T14:27:22.617Z). === 2025-07-04 === * 03:15 [[gitlab:galstar|@galstar]] was approved. === 2025-07-02 === * 11:27 "vicolas11" was rejected (pending since 2025-04-02T11:25:12.682Z). === 2025-06-29 === * 23:12 "naomi723" was rejected (pending since 2025-03-30T23:09:24.630Z). === 2025-06-28 === * 16:21 "mudeh2372" was rejected (pending since 2025-03-29T16:18:27.057Z). === 2025-06-27 === * 23:18 "rony143" was rejected (pending since 2025-03-28T23:16:13.671Z). * 22:21 [[gitlab:rluts|@rluts]] was approved. === 2025-06-26 === * 13:54 "creativegurus" was rejected (pending since 2025-03-27T13:52:41.706Z). === 2025-06-24 === * 17:42 [[gitlab:devjadiya|@devjadiya]] was approved. * 14:00 "dominic-r" was rejected (pending since 2025-03-25T14:00:07.307Z). === 2025-06-21 === * 00:48 [[gitlab:vriaa|@vriaa]] was approved. === 2025-06-18 === * 15:21 "ayushkhati1" was rejected (pending since 2025-03-19T15:18:50.062Z). === 2025-06-17 === * 20:45 "chiomavero" was rejected (pending since 2025-03-18T20:44:13.967Z). * 00:27 [[gitlab:eggroll97|@eggroll97]] was approved. === 2025-06-14 === * 20:57 "volvox" was rejected (pending since 2025-03-15T20:56:34.018Z). === 2025-06-13 === * 16:09 [[gitlab:supergrey|@supergrey]] was approved. * 11:03 "chqaz" was rejected (pending since 2025-03-14T11:01:09.600Z). * 10:24 [[gitlab:slong-wmf|@slong-wmf]] was approved. * 10:15 "hearvox" was rejected (pending since 2025-03-14T10:13:13.112Z). === 2025-06-12 === * 15:18 "jlam" was rejected (pending since 2025-03-13T15:17:54.099Z). === 2025-06-09 === * 20:48 "dipanjansengupta" was rejected (pending since 2025-03-10T20:48:03.545Z). * 19:27 [[gitlab:reggycelly|@reggycelly]] was approved. * 14:51 "arendpieter" was rejected (pending since 2025-03-10T14:51:01.445Z). * 13:21 [[gitlab:greenreaper|@greenreaper]] was approved. * 09:33 [[gitlab:mmta|@mmta]] was approved. * 08:03 "a-ssh22" was rejected (pending since 2025-03-10T08:03:08.111Z). === 2025-06-08 === * 21:06 "mm-episodenlistedlvaupdater" was rejected (pending since 2025-03-09T21:04:06.323Z). === 2025-06-06 === * 11:06 [[gitlab:olea|@olea]] was approved. === 2025-06-05 === * 20:33 [[gitlab:encodedwp|@encodedwp]] was approved. * 15:00 [[gitlab:toluayo|@toluayo]] was approved. * 13:51 [[gitlab:arnold_lup|@arnold_lup]] was approved. * 11:54 "sdhehua" was rejected (pending since 2025-03-06T11:51:48.241Z). === 2025-06-03 === * 21:27 [[gitlab:wewakey|@wewakey]] was approved. * 12:36 "hunsimon2" was rejected (pending since 2025-03-04T12:34:56.520Z). * 11:54 "hunsimon" was rejected (pending since 2025-03-04T11:53:54.652Z). === 2025-06-02 === * 12:01 [[gitlab:jaimedes|@jaimedes]] was approved. === 2025-05-30 === * 18:00 "sathvik9105" was rejected (pending since 2025-02-28T17:59:42.867Z). * 11:21 [[gitlab:tonythomas01|@tonythomas01]] was approved. * 10:06 [[gitlab:gpsleo|@gpsleo]] was approved. === 2025-05-29 === * 22:12 [[gitlab:codynguyen1116|@codynguyen1116]] was approved. === 2025-05-28 === * 02:57 [[gitlab:saper|@saper]] was approved. === 2025-05-27 === * 21:06 [[gitlab:mohammed_qays|@mohammed_qays]] was approved. * 15:33 "satanluimm" was rejected (pending since 2025-02-25T15:32:48.101Z). === 2025-05-26 === * 23:57 "seyedali220" was rejected (pending since 2025-02-24T23:56:17.621Z). === 2025-05-21 === * 11:12 [[gitlab:guilherme|@guilherme]] was approved. === 2025-05-19 === * 13:24 [[gitlab:emojiwiki|@emojiwiki]] was approved. === 2025-05-18 === * 00:00 "xidme" was rejected (pending since 2025-02-15T23:58:56.796Z). === 2025-05-17 === * 02:39 "kdh8219" was rejected (pending since 2025-02-15T02:36:32.237Z). === 2025-05-16 === * 15:09 [[gitlab:maxbinderwmf|@maxbinderwmf]] was approved. === 2025-05-15 === * 04:30 "inspectorzer0" was rejected (pending since 2025-02-13T04:27:33.179Z). === 2025-05-14 === * 17:42 [[gitlab:llugo|@llugo]] was approved. === 2025-05-13 === * 20:18 "mmta" was rejected (pending since 2025-02-11T20:17:23.407Z). === 2025-05-11 === * 20:51 "jad" was rejected (pending since 2025-02-09T20:49:07.333Z). * 17:54 "nishchalsundan" was rejected (pending since 2025-02-09T17:52:25.761Z). * 16:39 "mohammed_abukhadra" was rejected (pending since 2025-02-09T16:39:03.730Z). === 2025-05-09 === * 09:12 [[gitlab:sirchanmp|@sirchanmp]] was approved. === 2025-05-08 === * 08:18 [[gitlab:mengeditch|@mengeditch]] was approved. === 2025-05-07 === * 03:45 "xluffy" was rejected (pending since 2025-02-05T03:45:14.181Z). === 2025-05-06 === * 16:54 "punhaniabhishek" was rejected (pending since 2025-02-04T16:53:50.758Z). * 09:36 [[gitlab:bmartinezcalvo|@bmartinezcalvo]] was approved. === 2025-05-02 === * 12:24 [[gitlab:tohaomg|@tohaomg]] was approved. * 11:48 [[gitlab:mavrikant|@mavrikant]] was approved. * 11:45 [[gitlab:daanvr|@daanvr]] was approved. === 2025-05-01 === * 09:09 "mjoerg" was rejected (pending since 2025-01-30T09:09:04.204Z). === 2025-04-30 === * 23:06 "sanskardubey" was rejected (pending since 2025-01-29T23:03:25.489Z). === 2025-04-29 === * 16:00 "geyslein" was rejected (pending since 2025-01-28T16:00:01.510Z). === 2025-04-26 === * 09:30 "anjali9027" was rejected (pending since 2025-01-25T09:28:07.064Z). === 2025-04-25 === * 18:00 "salahhazaa" was rejected (pending since 2025-01-24T17:58:30.030Z). * 15:15 [[gitlab:yiming|@yiming]] was approved. * 02:06 "mrchanmp" was rejected (pending since 2025-01-24T02:03:58.308Z). === 2025-04-23 === * 17:03 "rj2904" was rejected (pending since 2025-01-22T17:03:11.207Z). * 14:21 "nischay33" was rejected (pending since 2025-01-22T14:19:21.081Z). === 2025-04-22 === * 19:27 "dj80" was rejected (pending since 2025-01-21T19:25:28.498Z). * 14:30 [[gitlab:kaimamin|@kaimamin]] was approved. * 09:57 "debo" was rejected (pending since 2025-01-21T09:54:47.955Z). === 2025-04-21 === * 12:24 "unshell" was rejected (pending since 2025-01-20T12:21:59.686Z). === 2025-04-18 === * 15:06 [[gitlab:spartanarbinger|@spartanarbinger]] was approved. === 2025-04-16 === * 03:09 "dewey" was rejected (pending since 2025-01-15T03:06:17.488Z). === 2025-04-15 === * 19:45 "emdadul" was rejected (pending since 2025-01-14T19:42:29.285Z). === 2025-04-14 === * 06:45 [[gitlab:bcampbell804|@bcampbell804]] was approved. === 2025-04-11 === * 06:27 [[gitlab:jvanderhoop|@jvanderhoop]] was approved. === 2025-04-10 === * 04:12 "bhai420" was rejected (pending since 2025-01-09T04:10:29.430Z). === 2025-04-09 === * 05:03 "austinvarshney" was rejected (pending since 2025-01-08T05:02:34.175Z). === 2025-04-06 === * 15:36 [[gitlab:elph|@elph]] was approved. === 2025-04-02 === * 10:33 [[gitlab:ozge|@ozge]] was approved. === 2025-03-31 === * 20:15 "demandkey" was rejected (pending since 2024-12-30T20:14:23.096Z). * 15:18 [[gitlab:danyya|@danyya]] was approved. === 2025-03-28 === * 15:54 [[gitlab:rutsavi09|@rutsavi09]] was approved. * 15:54 [[gitlab:ilanen1|@ilanen1]] was approved. === 2025-03-25 === * 19:27 [[gitlab:irfo|@irfo]] was approved. * 11:54 [[gitlab:kmontalva-wmf|@kmontalva-wmf]] was approved. * 04:33 [[gitlab:paul26|@paul26]] was approved. * 04:18 "as1100k" was rejected (pending since 2024-12-24T04:18:06.813Z). === 2025-03-24 === * 11:33 "amzadkhankk" was rejected (pending since 2024-12-23T11:33:14.176Z). === 2025-03-23 === * 12:24 "wolfdo" was rejected (pending since 2024-12-22T12:23:35.056Z). === 2025-03-22 === * 09:45 [[gitlab:fjmustak|@fjmustak]] was approved. === 2025-03-20 === * 18:42 "sathishkokila" was rejected (pending since 2024-12-19T18:39:35.161Z). * 17:03 [[gitlab:alien4444|@alien4444]] was approved. * 15:27 [[gitlab:davidcoronel|@davidcoronel]] was approved. === 2025-03-19 === * 22:57 [[gitlab:r1f4t|@r1f4t]] was approved. * 19:03 "daniel24ps" was rejected (pending since 2024-12-18T19:00:21.249Z). * 14:18 [[gitlab:beepbooppenguin|@beepbooppenguin]] was approved. === 2025-03-18 === * 17:48 "rahulkundu1209" was rejected (pending since 2024-12-17T17:46:41.936Z). * 08:15 "kirtisikka972" was rejected (pending since 2024-12-17T08:13:25.487Z). === 2025-03-15 === * 13:30 "tulspal_sidhu" was rejected (pending since 2024-12-14T13:29:10.606Z). * 01:39 "peacedeadc" was rejected (pending since 2024-12-14T01:37:36.579Z). === 2025-03-14 === * 03:51 [[gitlab:chuckthebuck|@chuckthebuck]] was approved. * 02:33 "yxngtrtxll" was rejected (pending since 2024-12-13T02:31:51.658Z). === 2025-03-13 === * 14:36 [[gitlab:iccander|@iccander]] was approved. === 2025-03-12 === * 23:21 "jokerchic36" was rejected (pending since 2024-12-11T23:21:00.670Z). * 15:30 [[gitlab:naomi|@naomi]] was approved. * 15:27 [[gitlab:cobi|@cobi]] was approved. === 2025-03-11 === * 12:42 "mohitvermaxx" was rejected (pending since 2024-12-10T12:40:56.967Z). === 2025-03-10 === * 16:51 [[gitlab:nanona15dobato|@nanona15dobato]] was approved. === 2025-03-09 === * 22:39 [[gitlab:jonkolbert|@jonkolbert]] was approved. * 20:45 [[gitlab:urbanecmtest2|@urbanecmtest2]] was approved. === 2025-03-07 === * 16:54 [[gitlab:hswan|@hswan]] was approved. * 14:42 [[gitlab:atitkov|@atitkov]] was approved. * 00:42 [[gitlab:infrastruktur|@infrastruktur]] was approved. === 2025-03-06 === * 17:21 "johnmann" was rejected (pending since 2024-12-05T17:19:24.995Z). === 2025-03-05 === * 07:33 [[gitlab:monx9494|@monx9494]] was approved. === 2025-03-02 === * 21:21 "paul26" was rejected (pending since 2024-12-01T21:20:19.681Z). === 2025-03-01 === * 19:15 [[gitlab:izno|@izno]] was approved. * 12:45 [[gitlab:nyerho|@nyerho]] was approved. === 2025-02-28 === * 18:27 [[gitlab:chuckonwumelu|@chuckonwumelu]] was approved. * 13:09 "ashwinpraveengo" was rejected (pending since 2024-11-29T13:07:47.240Z). * 00:18 "eduardoaugusto" was rejected (pending since 2024-11-29T00:17:43.372Z). === 2025-02-27 === * 20:39 "volkanurl" was rejected (pending since 2024-11-28T20:37:18.101Z). === 2025-02-24 === * 21:15 [[gitlab:feeglgeef|@feeglgeef]] was approved. * 20:18 [[gitlab:piaanalysis2|@piaanalysis2]] was approved. * 19:06 [[gitlab:dhardy|@dhardy]] was approved. === 2025-02-22 === * 19:27 [[gitlab:owuh|@owuh]] was approved. === 2025-02-19 === * 16:06 [[gitlab:artemkloko|@artemkloko]] was approved. * 13:03 [[gitlab:jgafnea|@jgafnea]] was approved. === 2025-02-17 === * 16:33 [[gitlab:asmartkitten|@asmartkitten]] was approved. === 2025-02-16 === * 19:12 "gaurigupta21" was rejected (pending since 2024-11-17T19:11:07.416Z). === 2025-02-15 === * 01:18 [[gitlab:mediawiki-quickstart-ci|@mediawiki-quickstart-ci]] was approved. === 2025-02-14 === * 15:21 "nathanbnm" was rejected (pending since 2024-11-15T15:18:19.632Z). === 2025-02-13 === * 16:45 [[gitlab:priyanshuchahal|@priyanshuchahal]] was approved. * 16:42 [[gitlab:ajhalili2006|@ajhalili2006]] was approved. === 2025-02-12 === * 23:21 "monkeypatch999" was rejected (pending since 2024-11-13T23:20:38.398Z). * 06:36 [[gitlab:jainlakshita28|@jainlakshita28]] was approved. === 2025-02-11 === * 19:27 [[gitlab:matthewsm2|@matthewsm2]] was approved. === 2025-02-09 === * 16:15 "mohammed_abukhadra" was rejected (pending since 2024-11-10T16:15:18.361Z). === 2025-02-07 === * 21:33 "brennan" was rejected (pending since 2024-11-08T21:31:07.351Z). === 2025-02-06 === * 08:24 "mmta" was rejected (pending since 2024-11-07T08:22:36.724Z). * 06:21 [[gitlab:bunnypranav|@bunnypranav]] was approved. === 2025-02-05 === * 22:39 "chrissteinchen" was rejected (pending since 2024-11-06T22:38:16.673Z). === 2025-02-03 === * 07:45 "edriiic" was rejected (pending since 2024-11-04T07:44:46.849Z). * 01:12 "geppy" was rejected (pending since 2024-11-04T01:10:48.710Z). === 2025-02-02 === * 13:18 "funa-enpitu" was rejected (pending since 2024-11-03T13:15:46.065Z). === 2025-01-31 === * 23:42 "nfontes" was rejected (pending since 2024-11-01T23:39:41.755Z). * 22:51 "sbronson" was rejected (pending since 2024-11-01T22:50:31.871Z). * 00:42 [[gitlab:farid|@farid]] was approved. === 2025-01-27 === * 08:15 [[gitlab:eliza189|@eliza189]] was approved. === 2025-01-25 === * 09:51 [[gitlab:pamputt|@pamputt]] was approved. === 2025-01-23 === * 14:30 [[gitlab:lubianat|@lubianat]] was approved. * 11:45 [[gitlab:bootsa|@bootsa]] was approved. === 2025-01-21 === * 05:09 "niko" was rejected (pending since 2024-07-21T16:10:01.377Z). * 05:09 "thawizkid369777" was rejected (pending since 2024-07-18T17:42:44.493Z). * 05:09 "sarthaksingh2" was rejected (pending since 2024-07-10T11:31:30.470Z). * 05:09 "shriyakt" was rejected (pending since 2024-07-06T04:54:10.248Z). * 05:09 "akshaya" was rejected (pending since 2024-07-06T04:04:51.488Z). * 05:09 "alaka03aj" was rejected (pending since 2024-07-05T18:01:54.876Z). * 05:09 "sulochanaviji-5049" was rejected (pending since 2024-07-01T05:58:00.427Z). * 05:09 "nayanjnath" was rejected (pending since 2024-07-01T02:51:57.405Z). * 05:09 "sd44" was rejected (pending since 2024-06-30T04:28:51.436Z). * 05:09 "metavalent" was rejected (pending since 2024-06-29T01:37:14.210Z). * 05:09 "wicloudx" was rejected (pending since 2024-06-28T11:51:23.335Z). * 05:09 "debo" was rejected (pending since 2024-06-28T01:44:59.845Z). * 05:09 "bwiki" was rejected (pending since 2024-06-23T14:15:38.032Z). * 05:09 "toprak" was rejected (pending since 2024-06-23T11:35:50.819Z). * 05:09 "iristeller" was rejected (pending since 2024-06-14T20:53:48.959Z). * 05:09 "jcolvin" was rejected (pending since 2024-06-12T17:29:01.238Z). * 05:09 "kalyan" was rejected (pending since 2024-06-07T07:52:46.993Z). * 05:09 "bluecrystal" was rejected (pending since 2024-06-06T19:16:20.107Z). * 05:09 "iftttrohit" was rejected (pending since 2024-06-04T12:08:50.818Z). * 05:09 "pogpotato" was rejected (pending since 2024-06-03T17:58:21.684Z). * 05:09 "cptlausebaer" was rejected (pending since 2024-05-31T18:53:27.692Z). * 05:09 "hdevine825" was rejected (pending since 2024-05-31T17:04:18.279Z). * 05:09 "anaghaa18" was rejected (pending since 2024-05-25T19:14:31.803Z). * 05:09 "atharvanair04" was rejected (pending since 2024-05-25T14:24:52.825Z). * 05:09 "anasvemmully" was rejected (pending since 2024-05-25T06:10:27.261Z). * 05:09 "abhinavmohandas" was rejected (pending since 2024-05-25T06:05:24.825Z). * 05:09 "kksurendran06" was rejected (pending since 2024-05-25T06:04:38.082Z). * 05:09 "albertmarshall8896" was rejected (pending since 2024-05-23T09:32:05.462Z). * 05:09 "akellison" was rejected (pending since 2024-05-17T02:07:24.229Z). * 05:09 "mainowill" was rejected (pending since 2024-04-16T23:30:33.881Z). * 05:09 "bzhqc" was rejected (pending since 2024-04-16T19:50:38.676Z). * 05:09 "safan41" was rejected (pending since 2024-04-16T03:34:48.942Z). * 05:09 "mgagat" was rejected (pending since 2024-04-16T03:21:51.764Z). * 05:09 "okeamah" was rejected (pending since 2024-04-16T02:49:00.143Z). * 05:09 "xuhao61" was rejected (pending since 2024-04-15T23:45:09.083Z). * 04:47 "cybel" was rejected (pending since 2024-04-15T06:46:35.791Z). === 2025-01-20 === * 14:33 [[gitlab:your1|@your1]] was approved. === 2025-01-18 === * 10:09 [[gitlab:galrach600|@galrach600]] was approved. * 02:51 [[gitlab:blankeclair|@blankeclair]] was approved. === 2025-01-17 === * 13:57 [[gitlab:dsantamaria|@dsantamaria]] was approved. === 2025-01-15 === * 17:12 [[gitlab:smartse|@smartse]] was approved. === 2025-01-14 === * 17:03 [[gitlab:naorleizer|@naorleizer]] was approved. === 2025-01-13 === * 02:45 [[gitlab:wolf20482|@wolf20482]] was approved. === 2025-01-12 === * 17:45 [[gitlab:tamzin|@tamzin]] was approved. === 2025-01-11 === * 15:24 [[gitlab:bargioni|@bargioni]] was approved. * 14:30 [[gitlab:salelya|@salelya]] was approved. * 10:15 [[gitlab:malakatshy|@malakatshy]] was approved. * 05:21 [[gitlab:newmcpee|@newmcpee]] was approved. === 2025-01-09 === * 15:30 [[gitlab:gkyziridis|@gkyziridis]] was approved. === 2025-01-08 === * 16:21 [[gitlab:ukrface|@ukrface]] was approved. === 2024-12-28 === * 03:27 [[gitlab:twonum|@twonum]] was approved. === 2024-12-25 === * 06:09 [[gitlab:harsv567|@harsv567]] was approved. === 2024-12-21 === * 11:24 [[gitlab:amutha2002|@amutha2002]] was approved. === 2024-12-20 === * 19:51 [[gitlab:hridyeshgupta|@hridyeshgupta]] was approved. * 10:00 [[gitlab:ro-shines|@ro-shines]] was approved. * 08:09 [[gitlab:kesharwaniarpita|@kesharwaniarpita]] was approved. === 2024-12-18 === * 14:45 [[gitlab:soylacarli|@soylacarli]] was approved. === 2024-12-16 === * 20:33 [[gitlab:aleyasiddika1|@aleyasiddika1]] was approved. === 2024-12-15 === * 07:33 [[gitlab:abhishek02bhardwaj|@abhishek02bhardwaj]] was approved. === 2024-12-13 === * 13:18 [[gitlab:ashmitabathre204|@ashmitabathre204]] was approved. === 2024-12-10 === * 06:39 [[gitlab:ginaan|@ginaan]] was approved. === 2024-12-09 === * 05:45 [[gitlab:kallinavya|@kallinavya]] was approved. * 00:54 [[gitlab:viserion-7|@viserion-7]] was approved. === 2024-12-08 === * 17:27 [[gitlab:wargo|@wargo]] was approved. === 2024-12-05 === * 11:15 [[gitlab:ranjithraj|@ranjithraj]] was approved. === 2024-12-02 === * 21:21 [[gitlab:a930913|@a930913]] was approved. === 2024-12-01 === * 02:39 [[gitlab:kingchristlike1|@kingchristlike1]] was approved. === 2024-11-21 === * 13:45 [[gitlab:sascha|@sascha]] was approved. === 2024-11-19 === * 16:36 [[gitlab:jly|@jly]] was approved. === 2024-11-15 === * 02:54 [[gitlab:danielyepezgarces|@danielyepezgarces]] was approved. === 2024-11-14 === * 14:15 [[gitlab:stimoroll|@stimoroll]] was approved. === 2024-11-09 === * 17:15 [[gitlab:f4udeveloper|@f4udeveloper]] was approved. === 2024-11-07 === * 19:15 [[gitlab:zulf|@zulf]] was approved. * 05:33 [[gitlab:hassanamin|@hassanamin]] was approved. === 2024-11-06 === * 19:39 [[gitlab:daniuu|@daniuu]] was approved. * 00:18 [[gitlab:rlopez-wmf|@rlopez-wmf]] was approved. === 2024-10-09 === * 14:45 [[gitlab:jtweed|@jtweed]] was approved. * 10:24 [[gitlab:ifrahkh|@ifrahkh]] was approved. * 09:06 [[gitlab:wikibayer|@wikibayer]] was approved. === 2024-10-06 === * 10:27 [[gitlab:keerthan16|@keerthan16]] was approved. === 2024-10-04 === * 07:45 [[gitlab:hakimi97|@hakimi97]] was approved. === 2024-09-30 === * 07:39 [[gitlab:ninjastrikers|@ninjastrikers]] was approved. === 2024-09-28 === * 17:30 [[gitlab:webrunner95|@webrunner95]] was approved. === 2024-09-18 === * 21:39 [[gitlab:elliottetzkorn|@elliottetzkorn]] was approved. === 2024-09-14 === * 22:06 [[gitlab:humptydumpty|@humptydumpty]] was approved. === 2024-09-06 === * 08:48 [[gitlab:mickabarber|@mickabarber]] was approved. === 2024-08-27 === * 17:36 [[gitlab:edgars|@edgars]] was approved. === 2024-08-22 === * 09:18 [[gitlab:antonkokhwmde|@antonkokhwmde]] was approved. === 2024-08-14 === * 19:21 [[gitlab:jfk|@jfk]] was approved. === 2024-08-13 === * 17:57 [[gitlab:daxserver|@daxserver]] was approved. === 2024-08-11 === * 09:57 [[gitlab:pauliesnug|@pauliesnug]] was approved. === 2024-08-10 === * 08:42 [[gitlab:ashig|@ashig]] was approved. === 2024-08-09 === * 14:09 [[gitlab:masssly|@masssly]] was approved. === 2024-08-05 === * 22:15 [[gitlab:mrtortue|@mrtortue]] was approved. === 2024-08-02 === * 16:21 [[gitlab:dsantini|@dsantini]] was approved. === 2024-07-31 === * 11:54 [[gitlab:cptviraj|@cptviraj]] was approved. === 2024-07-30 === * 19:09 [[gitlab:iniquity|@iniquity]] was approved. * 10:00 [[gitlab:collins|@collins]] was approved. === 2024-07-27 === * 15:57 [[gitlab:songnguxyz|@songnguxyz]] was approved. === 2024-07-25 === * 12:36 [[gitlab:mszabo|@mszabo]] was approved. * 09:21 [[gitlab:agarwalmahima|@agarwalmahima]] was approved. === 2024-07-24 === * 08:05 [[gitlab:dragoniez|@dragoniez]] was approved. === 2024-07-23 === * 06:54 [[gitlab:mirji|@mirji]] was approved. === 2024-07-16 === * 10:00 [[gitlab:lakejason0|@lakejason0]] was approved. === 2024-07-12 === * 11:33 [[gitlab:cn|@cn]] was approved. * 08:12 [[gitlab:unchampignon|@unchampignon]] was approved. === 2024-07-07 === * 17:12 [[gitlab:agamyasamuel|@agamyasamuel]] was approved. * 05:24 [[gitlab:kuldeepburjbhalaike|@kuldeepburjbhalaike]] was approved. === 2024-07-06 === * 11:18 [[gitlab:dibya|@dibya]] was approved. * 04:54 [[gitlab:sarthakparashar|@sarthakparashar]] was approved. === 2024-07-05 === * 18:15 [[gitlab:vanshikarathi|@vanshikarathi]] was approved. === 2024-07-02 === * 19:00 [[gitlab:ebrahim|@ebrahim]] was approved. === 2024-07-01 === * 20:12 [[gitlab:rockingpenny4|@rockingpenny4]] was approved. * 18:15 [[gitlab:balajijagadesh|@balajijagadesh]] was approved. === 2024-06-30 === * 18:24 [[gitlab:hrideshmg|@hrideshmg]] was approved. * 07:18 [[gitlab:chanakyakumardas|@chanakyakumardas]] was approved. * 06:30 [[gitlab:rihaan180|@rihaan180]] was approved. === 2024-06-27 === * 17:36 [[gitlab:driedmueller|@driedmueller]] was approved. === 2024-06-19 === * 12:57 [[gitlab:audreypenven|@audreypenven]] was approved. === 2024-06-16 === * 01:18 [[gitlab:roysmith|@roysmith]] was approved. === 2024-06-08 === * 02:45 [[gitlab:jleedev|@jleedev]] was approved. === 2024-06-03 === * 13:57 [[gitlab:afeder|@afeder]] was approved. === 2024-06-01 === * 10:54 [[gitlab:florianschmitt|@florianschmitt]] was approved. === 2024-05-30 === * 16:42 [[gitlab:krlsca|@krlsca]] was approved. === 2024-05-28 === * 11:24 [[gitlab:rickijay|@rickijay]] was approved. === 2024-05-26 === * 11:18 [[gitlab:ranjithsiji|@ranjithsiji]] was approved. === 2024-05-25 === * 07:24 [[gitlab:jony|@jony]] was approved. === 2024-05-23 === * 08:45 [[gitlab:lepticed7|@lepticed7]] was approved. === 2024-05-22 === * 20:42 [[gitlab:echecs|@echecs]] was approved. === 2024-05-21 === * 13:33 [[gitlab:mbs|@mbs]] was approved. === 2024-05-19 === * 18:06 [[gitlab:ionenlaser|@ionenlaser]] was approved. === 2024-05-18 === * 23:36 [[gitlab:mdaniels5757|@mdaniels5757]] was approved. === 2024-05-17 === * 08:54 [[gitlab:grapedog|@grapedog]] was approved. === 2024-05-08 === * 19:42 [[gitlab:kelhurd|@kelhurd]] was approved. * 19:06 [[gitlab:khurd|@khurd]] was approved. === 2024-05-06 === * 19:48 [[gitlab:j3j5|@j3j5]] was approved. * 12:06 [[gitlab:tk-999|@tk-999]] was approved. === 2024-05-05 === * 22:09 [[gitlab:pppery|@pppery]] was approved. * 20:33 [[gitlab:sakretsu|@sakretsu]] was approved. * 12:12 [[gitlab:waterquark|@waterquark]] was approved. === 2024-05-04 === * 09:03 [[gitlab:multichill|@multichill]] was approved. * 07:42 [[gitlab:abaris|@abaris]] was approved. === 2024-05-03 === * 14:57 [[gitlab:maurusian|@maurusian]] was approved. === 2024-04-24 === * 05:48 [[gitlab:wolfinux|@wolfinux]] was approved. === 2024-04-23 === * 15:48 [[gitlab:dreamrimmer|@dreamrimmer]] was approved. === 2024-04-21 === * 06:51 [[gitlab:alon|@alon]] was approved. === 2024-04-17 === * 23:33 [[gitlab:derenrich|@derenrich]] was approved. === 2024-04-16 === * 17:18 [[gitlab:valcio|@valcio]] was approved. === 2024-04-14 === * 16:51 [[gitlab:wikilucas00|@wikilucas00]] was approved. === 2024-04-06 === * 12:48 [[gitlab:theprotonade|@theprotonade]] was approved. === 2024-04-02 === * 07:30 [[gitlab:bohuizhang|@bohuizhang]] was approved. === 2024-03-30 === * 13:36 [[gitlab:lpintscher|@lpintscher]] was approved. === 2024-03-26 === * 17:09 [[gitlab:eenabulele|@eenabulele]] was approved. === 2024-03-25 === * 14:27 [[gitlab:tuukka|@tuukka]] was approved. === 2024-03-24 === * 12:24 [[gitlab:firefly|@firefly]] was approved. === 2024-03-21 === * 19:33 [[gitlab:universal-omega|@universal-omega]] was approved. === 2024-03-17 === * 10:36 [[gitlab:bisel91|@bisel91]] was approved. === 2024-03-16 === * 10:09 [[gitlab:delord|@delord]] was approved. * 00:42 [[gitlab:athulvis1|@athulvis1]] was approved. === 2024-03-15 === * 19:06 [[gitlab:ignaciorodrguez|@ignaciorodrguez]] was approved. * 08:30 [[gitlab:peachey88|@peachey88]] was approved. * 06:51 [[gitlab:derick|@derick]] was approved. === 2024-03-12 === * 15:06 [[gitlab:xiaoxiao|@xiaoxiao]] was approved. === 2024-03-06 === * 13:21 [[gitlab:desianabae1|@desianabae1]] was approved. === 2024-03-05 === * 19:21 [[gitlab:ep1c|@ep1c]] was approved. * 16:33 [[gitlab:jasmine|@jasmine]] was approved. === 2024-03-02 === * 06:42 [[gitlab:potsdamlamb|@potsdamlamb]] was approved. === 2024-02-29 === * 23:18 [[gitlab:arandomname123|@arandomname123]] was approved. * 18:03 [[gitlab:baba|@baba]] was approved. * 17:48 [[gitlab:yfdyh000|@yfdyh000]] was approved. * 03:09 [[gitlab:sds|@sds]] was approved. === 2024-02-27 === * 23:33 [[gitlab:lofhi|@lofhi]] was approved. === 2024-02-15 === * 19:45 [[gitlab:gergesshamon|@gergesshamon]] was approved. === 2024-02-14 === * 14:33 [[gitlab:philipnelson99|@philipnelson99]] was approved. === 2024-02-13 === * 13:06 [[gitlab:dringsim|@dringsim]] was approved. === 2024-02-12 === * 17:36 [[gitlab:haak|@haak]] was approved. === 2024-02-05 === * 17:33 [[gitlab:qwerfjkl|@qwerfjkl]] was approved. * 17:14 [[gitlab:ahecht|@ahecht]] was approved. === 2024-02-01 === * 09:27 [[gitlab:arinaigum|@arinaigum]] was approved. * 00:15 [[gitlab:jas42|@jas42]] was approved. * 00:15 [[gitlab:edhu|@edhu]] was approved. * 00:15 [[gitlab:marnanel|@marnanel]] was approved. * 00:15 [[gitlab:ibrahemqasim|@ibrahemqasim]] was approved. * 00:15 [[gitlab:amasotti|@amasotti]] was approved. * 00:15 [[gitlab:deni|@deni]] was approved. * 00:15 [[gitlab:cyber|@cyber]] was approved. * 00:15 [[gitlab:saroj|@saroj]] was approved. === 2024-01-29 === * 21:42 [[gitlab:rgupta|@rgupta]] was approved. === 2024-01-07 === * 09:48 [[gitlab:lutrome|@lutrome]] was approved. === 2024-01-05 === * 20:48 [[gitlab:jinoytommanjaly|@jinoytommanjaly]] was approved. * 02:51 [[gitlab:braunobruno|@braunobruno]] was approved. * 01:08 [[gitlab:amorymeltzer|@amorymeltzer]] was approved. * 01:08 [[gitlab:phi22ipus|@phi22ipus]] was approved. === 2024-01-03 === * 14:45 [[gitlab:gabina|@gabina]] was approved. === 2024-01-02 === * 13:18 [[gitlab:arthurtaylor|@arthurtaylor]] was approved. === 2023-12-23 === * 00:33 [[gitlab:aram|@aram]] was approved. === 2023-12-22 === * 16:24 [[gitlab:elpitareio|@elpitareio]] was approved. === 2023-12-21 === * 00:43 [[gitlab:bsadowski1|@bsadowski1]] was approved. * 00:43 [[gitlab:ederporto|@ederporto]] was approved. * 00:43 [[gitlab:sadraiiali|@sadraiiali]] was approved. * 00:43 [[gitlab:wasp-outis|@wasp-outis]] was approved. * 00:43 [[gitlab:bodhisattwa|@bodhisattwa]] was approved. * 00:43 [[gitlab:air7538|@air7538]] was approved. * 00:43 [[gitlab:anzx|@anzx]] was approved. * 00:43 [[gitlab:tekask1903|@tekask1903]] was approved. * 00:42 [[gitlab:kiwi-0x010c|@kiwi-0x010c]] was approved. * 00:42 [[gitlab:mpaa|@mpaa]] was approved. * 00:42 [[gitlab:kutay|@kutay]] was approved. * 00:42 [[gitlab:wattmto|@wattmto]] was approved. nuhslzq0nuk253q9xrrz9zy9yla2zcb 2445262 2445261 2026-08-09T21:18:26Z Gitlabaccountapprovalbot 37332 @iamnetx was approved. 2445262 wikitext text/x-wiki <noinclude>'''Audit log of approvals''' made by [[gitlab:gitlabaccountapprovalbot|@gitlabaccountapprovalbot]]. __NOTOC__</noinclude> === 2026-08-09 === * 21:18 [[gitlab:iamnetx|@iamnetx]] was approved. * 20:09 "yirba" was rejected (pending since 2026-05-10T20:07:07.738Z). * 06:42 "marsam2489" was rejected (pending since 2026-05-10T06:40:46.276Z). === 2026-08-08 === * 08:15 [[gitlab:taiwaniajusto|@taiwaniajusto]] was approved. === 2026-08-07 === * 18:21 [[gitlab:gturkington|@gturkington]] was approved. * 06:36 "brianbybyby" was rejected (pending since 2026-05-08T06:36:08.459Z). === 2026-07-30 === * 20:42 "horaciocolbert" was rejected (pending since 2026-04-30T20:41:40.421Z). === 2026-07-29 === * 10:03 [[gitlab:piastu|@piastu]] was approved. * 06:33 "rafiul1" was rejected (pending since 2026-04-29T06:33:02.573Z). === 2026-07-28 === * 16:51 [[gitlab:for-each-next|@for-each-next]] was approved. * 09:21 [[gitlab:gka|@gka]] was approved. === 2026-07-27 === * 08:12 [[gitlab:cambob|@cambob]] was approved. === 2026-07-26 === * 15:39 "demansanaagmailcom" was rejected (pending since 2026-04-26T15:39:03.297Z). === 2026-07-24 === * 10:54 [[gitlab:ysogo|@ysogo]] was approved. * 09:39 [[gitlab:pankaj199|@pankaj199]] was approved. === 2026-07-23 === * 12:30 [[gitlab:fermiboson|@fermiboson]] was approved. * 09:21 [[gitlab:slashme|@slashme]] was approved. === 2026-07-22 === * 15:54 [[gitlab:lmedley|@lmedley]] was approved. * 13:24 "praveen5638" was rejected (pending since 2026-04-22T13:21:23.368Z). * 12:48 [[gitlab:cyberpower678|@cyberpower678]] was approved. * 08:30 [[gitlab:nabbegat|@nabbegat]] was approved. * 08:30 [[gitlab:plyd|@plyd]] was approved. * 07:06 "ayush8620" was rejected (pending since 2026-04-22T07:03:12.476Z). === 2026-07-21 === * 15:51 [[gitlab:panieravide|@panieravide]] was approved. * 15:33 [[gitlab:luisvilla-personal|@luisvilla-personal]] was approved. * 13:48 [[gitlab:yru|@yru]] was approved. * 13:09 [[gitlab:deevad|@deevad]] was approved. * 12:54 [[gitlab:ctdo17|@ctdo17]] was approved. * 12:36 [[gitlab:jeannenoiraud|@jeannenoiraud]] was approved. * 12:21 [[gitlab:nadiantara|@nadiantara]] was approved. * 12:18 [[gitlab:wijltcher|@wijltcher]] was approved. * 10:33 [[gitlab:nivopol|@nivopol]] was approved. * 10:30 [[gitlab:johlig|@johlig]] was approved. * 10:30 [[gitlab:majicita|@majicita]] was approved. * 10:15 [[gitlab:yongjiapeng|@yongjiapeng]] was approved. * 10:09 [[gitlab:francyskus|@francyskus]] was approved. * 09:24 [[gitlab:xanonymusx|@xanonymusx]] was approved. === 2026-07-20 === * 17:15 "leonidlednev" was rejected (pending since 2026-04-20T17:13:35.108Z). * 15:27 [[gitlab:rodrigoargenton|@rodrigoargenton]] was approved. * 05:48 "draftecho" was rejected (pending since 2026-04-20T05:48:06.953Z). === 2026-07-19 === * 12:21 [[gitlab:boivie|@boivie]] was approved. === 2026-07-18 === * 16:09 [[gitlab:pharos|@pharos]] was approved. * 15:45 [[gitlab:priyankar22|@priyankar22]] was approved. * 15:30 [[gitlab:sisyph|@sisyph]] was approved. === 2026-07-13 === * 03:45 [[gitlab:dreamyshade|@dreamyshade]] was approved. === 2026-07-12 === * 09:27 [[gitlab:smk|@smk]] was approved. === 2026-07-11 === * 14:48 "bigcereal42" was rejected (pending since 2026-04-11T14:47:40.321Z). * 12:18 "pratyushsawan" was rejected (pending since 2026-04-11T12:16:41.671Z). === 2026-07-10 === * 14:42 [[gitlab:akaza24|@akaza24]] was approved. * 12:57 [[gitlab:kormisk|@kormisk]] was approved. === 2026-07-07 === * 10:36 [[gitlab:olafjanssen|@olafjanssen]] was approved. * 06:57 "elisapoly-99" was rejected (pending since 2026-04-07T06:55:52.662Z). === 2026-07-06 === * 11:57 "ma3rouf" was rejected (pending since 2026-04-06T11:56:30.978Z). === 2026-07-02 === * 15:48 [[gitlab:tekneos|@tekneos]] was approved. === 2026-07-01 === * 15:12 [[gitlab:mugurolevy|@mugurolevy]] was approved. * 14:15 [[gitlab:vadymts1|@vadymts1]] was approved. * 09:57 "mugurolevy" was rejected (pending since 2026-04-01T09:55:19.175Z). === 2026-06-30 === * 14:27 "shivangisharma" was rejected (pending since 2026-03-31T14:26:44.932Z). === 2026-06-29 === * 19:03 [[gitlab:thisismattmiller|@thisismattmiller]] was approved. === 2026-06-28 === * 14:51 "nkwenuinadine" was rejected (pending since 2026-03-29T14:48:32.735Z). * 14:03 "vaishnavikumbhar" was rejected (pending since 2026-03-29T14:01:30.604Z). * 13:03 "stepmay" was rejected (pending since 2026-03-29T13:01:59.905Z). * 06:51 "swallroth" was rejected (pending since 2026-03-29T06:49:54.838Z). === 2026-06-26 === * 09:39 [[gitlab:lakshita28|@lakshita28]] was approved. * 07:30 [[gitlab:reeti|@reeti]] was approved. * 07:30 [[gitlab:anushka10patel|@anushka10patel]] was approved. * 07:30 "samsaesque" was rejected (pending since 2026-03-27T07:29:57.279Z). * 05:51 [[gitlab:arpithhhaaa|@arpithhhaaa]] was approved. * 05:51 [[gitlab:govindlaltl|@govindlaltl]] was approved. === 2026-06-25 === * 16:09 [[gitlab:sakuraemad|@sakuraemad]] was approved. * 07:00 "kdh8219" was rejected (pending since 2026-03-26T06:58:05.415Z). === 2026-06-24 === * 11:54 [[gitlab:sanskardubeydev|@sanskardubeydev]] was approved. * 10:09 "tanmay789q" was rejected (pending since 2026-03-25T10:07:54.602Z). === 2026-06-22 === * 19:57 [[gitlab:gouvernathor|@gouvernathor]] was approved. * 16:45 [[gitlab:lucasbelo|@lucasbelo]] was approved. * 07:15 "jason2000-cpu" was rejected (pending since 2026-03-23T07:14:09.184Z). === 2026-06-21 === * 13:18 [[gitlab:egonw|@egonw]] was approved. === 2026-06-20 === * 10:21 [[gitlab:tways2017|@tways2017]] was approved. === 2026-06-19 === * 16:06 "wilsonwang2026" was rejected (pending since 2026-03-20T16:06:05.511Z). * 04:12 [[gitlab:claudio|@claudio]] was approved. === 2026-06-18 === * 14:21 "royiswariii" was rejected (pending since 2026-03-19T14:19:16.896Z). * 13:06 [[gitlab:laurabarluzzi|@laurabarluzzi]] was approved. === 2026-06-17 === * 11:24 "adinathq8x" was rejected (pending since 2026-03-18T11:22:50.098Z). * 09:45 "nathanveritas" was rejected (pending since 2026-03-18T09:43:51.645Z). === 2026-06-15 === * 22:39 [[gitlab:mohammadhijjawi|@mohammadhijjawi]] was approved. * 14:24 "enlisar" was rejected (pending since 2026-03-16T14:23:00.109Z). * 14:06 "ayaan" was rejected (pending since 2026-03-16T14:03:31.071Z). * 10:54 "kwametech" was rejected (pending since 2026-03-16T10:54:11.083Z). === 2026-06-14 === * 17:45 [[gitlab:surajseth520|@surajseth520]] was approved. * 07:24 "malahimhaseeb" was rejected (pending since 2026-03-15T07:21:57.748Z). === 2026-06-11 === * 11:48 [[gitlab:cadddr|@cadddr]] was approved. * 11:18 "wikipiggy" was rejected (pending since 2026-03-12T11:16:09.335Z). * 07:15 [[gitlab:vesihiisi|@vesihiisi]] was approved. === 2026-06-10 === * 07:03 [[gitlab:dmiranda|@dmiranda]] was approved. === 2026-06-09 === * 14:21 [[gitlab:linkgenetic|@linkgenetic]] was approved. * 14:03 [[gitlab:sjones-ctr|@sjones-ctr]] was approved. * 12:51 [[gitlab:ekrem|@ekrem]] was approved. === 2026-06-08 === * 17:48 "jmprax" was rejected (pending since 2026-03-09T17:46:38.807Z). * 16:15 [[gitlab:ahonc|@ahonc]] was approved. * 12:57 [[gitlab:rainmonger|@rainmonger]] was approved. === 2026-06-07 === * 23:03 "shadowthewuff" was rejected (pending since 2026-03-08T23:00:53.442Z). * 11:45 "wiki-pavan" was rejected (pending since 2026-03-08T11:45:11.116Z). * 02:39 [[gitlab:launchpad|@launchpad]] was approved. === 2026-06-06 === * 14:54 "unicord" was rejected (pending since 2026-03-07T14:52:04.992Z). * 12:48 "chien" was rejected (pending since 2026-03-07T12:48:11.669Z). === 2026-06-04 === * 14:33 "only-vikas" was rejected (pending since 2026-03-05T14:32:09.186Z). === 2026-06-03 === * 15:00 [[gitlab:anafibnshahibul|@anafibnshahibul]] was approved. === 2026-06-02 === * 21:21 "mgagat" was rejected (pending since 2026-03-03T21:18:37.223Z). * 13:57 "prasunaenumarthy" was rejected (pending since 2026-03-03T13:57:14.847Z). * 05:48 [[gitlab:tmoney|@tmoney]] was approved. === 2026-06-01 === * 14:57 "vikram2101" was rejected (pending since 2026-03-02T14:54:26.550Z). * 12:03 "watshell" was rejected (pending since 2026-03-02T12:03:09.329Z). === 2026-05-29 === * 12:48 "mounikapotladurthi" was rejected (pending since 2026-02-27T12:45:38.609Z). === 2026-05-27 === * 20:00 "vinitha" was rejected (pending since 2026-02-25T19:58:43.524Z). * 16:30 "codeurluce" was rejected (pending since 2026-02-25T16:28:53.973Z). * 14:33 [[gitlab:thilio|@thilio]] was approved. === 2026-05-26 === * 12:09 "charisad" was rejected (pending since 2026-02-24T12:07:21.881Z). === 2026-05-25 === * 22:54 "ddshelto" was rejected (pending since 2026-02-23T22:52:44.427Z). * 19:51 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z). * 19:48 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z). === 2026-05-24 === * 18:45 "jiyagupta-cs" was rejected (pending since 2026-02-22T18:43:33.176Z). === 2026-05-23 === * 13:09 [[gitlab:gauthammohanraj|@gauthammohanraj]] was approved. * 04:21 [[gitlab:staraction|@staraction]] was approved. === 2026-05-22 === * 19:03 "i-horich" was rejected (pending since 2026-02-20T19:00:43.519Z). * 01:48 "50323233" was rejected (pending since 2026-02-20T01:48:05.555Z). === 2026-05-21 === * 18:51 "kartikeyg0104" was rejected (pending since 2026-02-19T18:48:39.707Z). * 16:27 [[gitlab:renovatebot|@renovatebot]] was approved. * 16:06 [[gitlab:gkm563|@gkm563]] was approved. === 2026-05-20 === * 01:21 "beedellrokejulianlockhart" was rejected (pending since 2026-02-18T01:19:13.284Z). === 2026-05-18 === * 23:18 "wladek92" was rejected (pending since 2026-02-16T23:16:22.939Z). * 16:36 [[gitlab:effeietsanders|@effeietsanders]] was approved. === 2026-05-14 === * 21:00 [[gitlab:nehemienathan|@nehemienathan]] was approved. === 2026-05-13 === * 10:51 "ssssaaaa" was rejected (pending since 2026-02-11T10:50:36.975Z). === 2026-05-12 === * 18:06 [[gitlab:psubhashish|@psubhashish]] was approved. * 08:12 "khan" was rejected (pending since 2026-02-10T08:11:48.776Z). * 04:27 "galaxysh" was rejected (pending since 2026-02-10T04:24:59.440Z). === 2026-05-11 === * 12:18 "peterxy12" was rejected (pending since 2026-02-09T12:18:01.982Z). === 2026-05-10 === * 11:09 "yalihupokn" was rejected (pending since 2026-02-08T11:06:51.336Z). * 05:12 "wobadha" was rejected (pending since 2026-02-08T05:11:00.569Z). === 2026-05-09 === * 13:45 "bwiki" was rejected (pending since 2026-02-07T13:43:38.177Z). === 2026-05-08 === * 09:24 [[gitlab:cwilliams|@cwilliams]] was approved. === 2026-05-07 === * 14:15 "rehankhan78" was rejected (pending since 2026-02-05T14:13:37.754Z). === 2026-05-06 === * 11:24 "ari" was rejected (pending since 2026-02-04T11:24:11.760Z). * 08:09 [[gitlab:neriah|@neriah]] was approved. * 06:27 [[gitlab:status401|@status401]] was approved. === 2026-05-03 === * 09:54 [[gitlab:anilk|@anilk]] was approved. === 2026-05-02 === * 17:54 [[gitlab:sweil|@sweil]] was approved. * 17:00 [[gitlab:aoppo|@aoppo]] was approved. === 2026-05-01 === * 21:18 [[gitlab:dawalda|@dawalda]] was approved. === 2026-04-30 === * 21:42 "merohibine" was rejected (pending since 2026-01-29T21:40:00.756Z). * 20:54 [[gitlab:tfmorris|@tfmorris]] was approved. * 17:33 [[gitlab:uyen|@uyen]] was approved. * 07:39 [[gitlab:mahveotm|@mahveotm]] was approved. * 06:36 [[gitlab:leo321|@leo321]] was approved. === 2026-04-29 === * 02:27 [[gitlab:dw31415|@dw31415]] was approved. === 2026-04-28 === * 23:09 [[gitlab:dtorsani|@dtorsani]] was approved. === 2026-04-27 === * 23:42 [[gitlab:quinlan|@quinlan]] was approved. * 05:00 [[gitlab:matthewyeager|@matthewyeager]] was approved. === 2026-04-26 === * 17:36 "kuba-hajnej" was rejected (pending since 2026-01-25T17:33:32.467Z). * 13:03 "jklamo" was rejected (pending since 2026-01-25T13:02:22.936Z). === 2026-04-25 === * 20:24 [[gitlab:maldaxura|@maldaxura]] was approved. * 14:33 [[gitlab:sirtobi|@sirtobi]] was approved. * 04:18 "ice5678" was rejected (pending since 2026-01-24T04:15:30.008Z). === 2026-04-24 === * 22:06 [[gitlab:arcstur|@arcstur]] was approved. === 2026-04-22 === * 23:06 "dtorsani" was rejected (pending since 2026-01-21T23:03:25.843Z). * 22:18 [[gitlab:egezort|@egezort]] was approved. * 16:45 "nexpectarpit" was rejected (pending since 2026-01-21T16:43:21.045Z). === 2026-04-20 === * 19:15 "fitch" was rejected (pending since 2026-01-19T19:12:35.644Z). === 2026-04-19 === * 02:54 [[gitlab:neoact|@neoact]] was approved. === 2026-04-18 === * 07:06 [[gitlab:kockaadmiralac|@kockaadmiralac]] was approved. === 2026-04-17 === * 13:42 "liselot" was rejected (pending since 2026-01-16T13:39:41.909Z). === 2026-04-15 === * 17:03 "lahari" was rejected (pending since 2026-01-14T17:02:06.275Z). === 2026-04-14 === * 13:00 "surajseth520" was rejected (pending since 2026-01-13T12:59:45.906Z). * 04:51 [[gitlab:canley|@canley]] was approved. * 01:03 "bshizzle" was rejected (pending since 2026-01-13T01:00:48.120Z). === 2026-04-13 === * 15:30 [[gitlab:passimacopoulos|@passimacopoulos]] was approved. === 2026-04-11 === * 12:30 "krithash" was rejected (pending since 2026-01-10T12:27:24.731Z). === 2026-04-10 === * 15:30 "raunak1709" was rejected (pending since 2026-01-09T15:29:10.901Z). === 2026-04-07 === * 17:03 [[gitlab:supnabla|@supnabla]] was approved. === 2026-04-06 === * 20:00 [[gitlab:laerdon|@laerdon]] was approved. * 19:21 [[gitlab:ljq3|@ljq3]] was approved. === 2026-04-04 === * 11:06 "mixcc" was rejected (pending since 2026-01-03T11:03:33.922Z). === 2026-04-02 === * 05:30 [[gitlab:mbh1|@mbh1]] was approved. === 2026-04-01 === * 18:21 "yuvrajpatil17" was rejected (pending since 2025-12-31T18:20:27.991Z). * 12:12 [[gitlab:amorii0|@amorii0]] was approved. === 2026-03-31 === * 11:00 "krrishsehgal" was rejected (pending since 2025-12-30T11:00:16.384Z). === 2026-03-30 === * 15:36 [[gitlab:atsuko|@atsuko]] was approved. === 2026-03-29 === * 11:36 [[gitlab:giftcup|@giftcup]] was approved. === 2026-03-28 === * 14:51 [[gitlab:janeeva1|@janeeva1]] was approved. === 2026-03-26 === * 13:36 [[gitlab:saiphani02|@saiphani02]] was approved. * 11:48 [[gitlab:valerioboz-wmch|@valerioboz-wmch]] was approved. === 2026-03-25 === * 09:45 "quansi" was rejected (pending since 2025-12-24T09:42:13.451Z). * 02:18 [[gitlab:viztor|@viztor]] was approved. === 2026-03-24 === * 23:18 [[gitlab:maryyann|@maryyann]] was approved. * 23:01 [[gitlab:codenamenoreste|@codenamenoreste]] was approved. * 13:36 [[gitlab:marc-maillard-wmse|@marc-maillard-wmse]] was approved. * 07:39 "fred2675" was rejected (pending since 2025-12-23T07:39:11.380Z). === 2026-03-23 === * 14:51 [[gitlab:komla|@komla]] was approved. * 05:51 "lunachuck43" was rejected (pending since 2025-12-22T05:50:17.862Z). * 04:06 "reza110011" was rejected (pending since 2025-12-22T04:05:25.117Z). === 2026-03-20 === * 21:54 "mertgor" was rejected (pending since 2025-12-19T21:51:51.419Z). * 20:57 "autanmahmah" was rejected (pending since 2025-12-19T20:54:51.678Z). * 09:57 [[gitlab:nethahussain|@nethahussain]] was approved. * 09:27 [[gitlab:piewriter|@piewriter]] was approved. * 08:15 [[gitlab:dondersmooi|@dondersmooi]] was approved. === 2026-03-19 === * 21:03 "sayvhior" was rejected (pending since 2025-12-18T21:02:31.699Z). === 2026-03-18 === * 20:15 [[gitlab:martinmystere|@martinmystere]] was approved. === 2026-03-17 === * 02:51 "louperivois" was rejected (pending since 2025-12-16T02:50:48.197Z). === 2026-03-16 === * 12:54 "mokayaj857" was rejected (pending since 2025-12-15T12:53:39.015Z). * 06:18 "roamer15" was rejected (pending since 2025-12-15T06:16:38.042Z). === 2026-03-14 === * 11:12 "umaramuhammad" was rejected (pending since 2025-12-13T11:10:44.004Z). * 09:33 "akuma19" was rejected (pending since 2025-12-13T09:31:39.044Z). * 07:06 [[gitlab:syunsyunminmin|@syunsyunminmin]] was approved. === 2026-03-12 === * 20:24 [[gitlab:11wb|@11wb]] was approved. * 09:54 [[gitlab:bcxfu75k|@bcxfu75k]] was approved. === 2026-03-10 === * 09:12 [[gitlab:viktoriahillerudwmse|@viktoriahillerudwmse]] was approved. === 2026-03-06 === * 08:09 "vazhayilnewone" was rejected (pending since 2025-12-05T08:07:02.184Z). === 2026-03-04 === * 20:54 [[gitlab:elphie|@elphie]] was approved. * 11:39 "ronaldahmed" was rejected (pending since 2025-12-03T11:37:47.492Z). * 02:12 "ltslw" was rejected (pending since 2025-12-03T02:11:52.040Z). === 2026-03-02 === * 19:21 "dlopez350" was rejected (pending since 2025-12-01T19:20:38.918Z). * 18:15 [[gitlab:lsandergreen|@lsandergreen]] was approved. === 2026-03-01 === * 10:51 [[gitlab:clintacc|@clintacc]] was approved. === 2026-02-28 === * 09:24 "cardboardlamp" was rejected (pending since 2025-11-29T09:22:03.947Z). * 08:18 "wiki-pavan" was rejected (pending since 2025-11-29T08:16:24.184Z). === 2026-02-27 === * 20:45 "thisisrick25" was rejected (pending since 2025-11-28T20:42:24.454Z). === 2026-02-26 === * 13:57 "chuiimuiiofc" was rejected (pending since 2025-11-27T13:57:02.794Z). * 13:54 "steffpro" was rejected (pending since 2025-11-27T13:52:10.859Z). === 2026-02-25 === * 21:24 "abubakarhabibudayyabu" was rejected (pending since 2025-11-26T21:22:37.776Z). === 2026-02-24 === * 05:00 "playboi" was rejected (pending since 2025-11-25T05:00:30.762Z). === 2026-02-23 === * 14:00 "alph65" was rejected (pending since 2025-11-24T13:59:00.797Z). * 12:33 [[gitlab:robertsky|@robertsky]] was approved. === 2026-02-22 === * 00:30 "hp8p" was rejected (pending since 2025-11-23T00:29:24.741Z). === 2026-02-19 === * 16:45 "clayjar" was rejected (pending since 2025-11-20T16:44:48.380Z). === 2026-02-18 === * 22:18 "nexus" was rejected (pending since 2025-11-19T22:16:48.818Z). * 12:00 "bernsteinnn" was rejected (pending since 2025-11-19T11:59:04.427Z). === 2026-02-17 === * 11:36 "jason2000-cpu" was rejected (pending since 2025-11-18T11:34:00.314Z). === 2026-02-16 === * 14:54 "smaurya" was rejected (pending since 2025-11-17T14:52:06.906Z). === 2026-02-15 === * 16:51 "kra-79" was rejected (pending since 2025-11-16T16:50:41.375Z). === 2026-02-14 === * 15:15 [[gitlab:mess|@mess]] was approved. === 2026-02-13 === * 13:57 "sopalsuemae957" was rejected (pending since 2025-11-14T13:55:16.921Z). * 13:30 [[gitlab:wyslijp16-toolforge|@wyslijp16-toolforge]] was approved. === 2026-02-12 === * 16:30 "kristinagligoric" was rejected (pending since 2025-11-13T16:29:21.646Z). * 03:33 [[gitlab:anyehansen|@anyehansen]] was approved. * 02:21 [[gitlab:thejoyfultentmaker|@thejoyfultentmaker]] was approved. === 2026-02-10 === * 13:18 [[gitlab:db111|@db111]] was approved. === 2026-02-09 === * 19:06 "squirrel289" was rejected (pending since 2025-11-10T19:04:27.831Z). === 2026-02-06 === * 20:54 [[gitlab:gillux|@gillux]] was approved. * 09:09 [[gitlab:lih|@lih]] was approved. === 2026-01-31 === * 16:21 [[gitlab:taxonbot1|@taxonbot1]] was approved. === 2026-01-28 === * 14:30 [[gitlab:ademola|@ademola]] was approved. * 10:51 "watshell" was rejected (pending since 2025-10-29T10:51:01.521Z). === 2026-01-26 === * 23:06 "tavaresgmg" was rejected (pending since 2025-10-27T23:04:42.140Z). === 2026-01-25 === * 06:03 "cata" was rejected (pending since 2025-10-26T06:01:26.155Z). === 2026-01-24 === * 21:15 [[gitlab:wiegels|@wiegels]] was approved. * 06:30 [[gitlab:blaquans|@blaquans]] was approved. === 2026-01-23 === * 16:27 [[gitlab:lerickson|@lerickson]] was approved. * 10:15 "fran0035g" was rejected (pending since 2025-10-24T10:12:17.732Z). === 2026-01-22 === * 21:00 "hacksyn" was rejected (pending since 2025-10-23T20:59:15.982Z). === 2026-01-21 === * 17:30 [[gitlab:otcenas11|@otcenas11]] was approved. === 2026-01-19 === * 21:48 [[gitlab:amdrel|@amdrel]] was approved. * 04:36 "rayalexa" was rejected (pending since 2025-10-20T04:35:02.094Z). === 2026-01-18 === * 15:45 "somya" was rejected (pending since 2025-10-19T15:43:43.701Z). * 06:54 "sergg001" was rejected (pending since 2025-10-19T06:54:12.296Z). === 2026-01-16 === * 11:57 "zeejohsy" was rejected (pending since 2025-10-17T11:56:22.372Z). * 04:45 "rocky25" was rejected (pending since 2025-10-17T04:43:33.180Z). === 2026-01-15 === * 16:39 "tiisu" was rejected (pending since 2025-10-16T16:37:18.438Z). * 12:00 "noahalorwu" was rejected (pending since 2025-10-16T11:58:26.133Z). * 10:39 "prjayaiuedu" was rejected (pending since 2025-10-16T10:37:16.947Z). === 2026-01-13 === * 17:21 [[gitlab:lwilson-ctr|@lwilson-ctr]] was approved. === 2026-01-12 === * 17:03 "stagietechs" was rejected (pending since 2025-10-13T17:02:25.281Z). === 2026-01-10 === * 19:06 "keerthisr" was rejected (pending since 2025-10-11T19:05:01.758Z). === 2026-01-09 === * 20:36 "lightb" was rejected (pending since 2025-10-10T20:34:20.264Z). === 2026-01-08 === * 19:42 [[gitlab:tbodt|@tbodt]] was approved. * 13:57 [[gitlab:martynranyard|@martynranyard]] was approved. === 2026-01-07 === * 17:48 [[gitlab:santanuwiki25|@santanuwiki25]] was approved. * 14:27 "dipanshu" was rejected (pending since 2025-10-08T14:26:10.794Z). * 12:30 "adeolaadesina" was rejected (pending since 2025-10-08T12:29:49.592Z). * 09:21 "tony-kamande" was rejected (pending since 2025-10-08T09:20:28.421Z). * 06:18 "hninwuttyi" was rejected (pending since 2025-10-08T06:17:28.006Z). * 05:09 "andume" was rejected (pending since 2025-10-08T05:07:18.582Z). * 02:00 "mosope" was rejected (pending since 2025-10-08T01:59:54.800Z). * 01:15 [[gitlab:tungstalite|@tungstalite]] was approved. === 2026-01-06 === * 18:24 "leerensucher" was rejected (pending since 2025-10-07T18:21:41.253Z). * 14:54 "leonidlednev" was rejected (pending since 2025-10-07T14:53:07.273Z). * 12:57 "alexandre-tingaud" was rejected (pending since 2025-10-07T12:54:27.206Z). === 2026-01-04 === * 21:33 [[gitlab:matr1x-101|@matr1x-101]] was approved. * 15:18 "makjr" was rejected (pending since 2025-10-05T15:16:31.558Z). * 14:09 "dakshq" was rejected (pending since 2025-10-05T14:08:40.608Z). === 2026-01-03 === * 20:42 [[gitlab:apehitkey|@apehitkey]] was approved. * 18:00 [[gitlab:jeremyb|@jeremyb]] was approved. * 14:09 [[gitlab:twelephant|@twelephant]] was approved. === 2026-01-01 === * 11:30 "shellstanislav" was rejected (pending since 2025-10-02T11:29:10.150Z). === 2025-12-30 === * 19:51 "camilojdiaz" was rejected (pending since 2025-09-30T19:49:24.913Z). === 2025-12-29 === * 16:03 "zied" was rejected (pending since 2025-09-29T16:01:30.415Z). * 08:18 "rahulsidpradhan" was rejected (pending since 2025-09-29T08:17:02.849Z). === 2025-12-26 === * 09:48 "thembo42" was rejected (pending since 2025-09-26T09:45:15.033Z). === 2025-12-25 === * 14:03 "196936074751" was rejected (pending since 2025-09-25T14:02:31.367Z). === 2025-12-23 === * 16:21 "ngarnsworthy" was rejected (pending since 2025-09-23T16:20:41.211Z). === 2025-12-22 === * 12:39 "aza555" was rejected (pending since 2025-09-22T12:38:02.622Z). === 2025-12-20 === * 23:45 "saph" was rejected (pending since 2025-09-20T23:45:01.222Z). === 2025-12-19 === * 10:15 "vladdymoses" was rejected (pending since 2025-09-19T10:15:00.999Z). * 07:15 "dirtylittlepoobah" was rejected (pending since 2025-09-19T07:13:55.537Z). === 2025-12-18 === * 16:24 [[gitlab:guyfawcus|@guyfawcus]] was approved. === 2025-12-17 === * 21:39 [[gitlab:holdyourhorses|@holdyourhorses]] was approved. * 18:30 "prudencia" was rejected (pending since 2025-09-17T18:27:18.860Z). * 02:24 "lottie" was rejected (pending since 2025-09-17T02:21:21.744Z). === 2025-12-16 === * 09:39 [[gitlab:melcatherine|@melcatherine]] was approved. * 08:54 [[gitlab:leila237|@leila237]] was approved. === 2025-12-15 === * 18:27 [[gitlab:royalsailor|@royalsailor]] was approved. * 09:39 [[gitlab:olaf8940|@olaf8940]] was approved. * 09:39 "brianbybyby" was rejected (pending since 2025-09-15T09:37:45.430Z). === 2025-12-14 === * 20:21 [[gitlab:essa237|@essa237]] was approved. * 16:42 [[gitlab:bovimacoco|@bovimacoco]] was approved. === 2025-12-13 === * 21:54 "mmns21" was rejected (pending since 2025-09-13T21:52:24.017Z). * 20:33 "bugcrawler" was rejected (pending since 2025-09-13T20:31:09.211Z). === 2025-12-12 === * 14:39 "ruvchoudhary" was rejected (pending since 2025-09-12T14:36:16.167Z). * 06:54 "rezadress" was rejected (pending since 2025-09-12T06:52:21.749Z). === 2025-12-10 === * 17:30 [[gitlab:itsmoon|@itsmoon]] was approved. === 2025-12-09 === * 15:42 [[gitlab:mercy-o|@mercy-o]] was approved. === 2025-12-06 === * 16:45 "jacquesradjabu" was rejected (pending since 2025-09-06T16:45:17.969Z). * 11:27 [[gitlab:ikhitron|@ikhitron]] was approved. === 2025-12-01 === * 08:12 "halconmilenario21" was rejected (pending since 2025-09-01T08:12:10.262Z). === 2025-11-30 === * 21:06 [[gitlab:habs|@habs]] was approved. === 2025-11-29 === * 16:36 "bovimacoco" was rejected (pending since 2025-08-30T16:34:39.712Z). * 00:45 [[gitlab:jjpmaster|@jjpmaster]] was approved. === 2025-11-24 === * 10:30 "alph65" was rejected (pending since 2025-08-25T10:28:40.957Z). * 02:24 [[gitlab:yaron|@yaron]] was approved. === 2025-11-20 === * 16:06 "clayjar" was rejected (pending since 2025-08-21T16:04:54.450Z). === 2025-11-17 === * 21:09 [[gitlab:ankita97531|@ankita97531]] was approved. === 2025-11-16 === * 14:15 "commanderkefir" was rejected (pending since 2025-08-17T14:13:14.791Z). * 08:21 "rehankhan78" was rejected (pending since 2025-08-17T08:19:44.896Z). === 2025-11-15 === * 14:36 "cyberscribe" was rejected (pending since 2025-08-16T14:34:27.230Z). === 2025-11-13 === * 04:21 "waddie96" was rejected (pending since 2025-08-14T04:19:27.461Z). === 2025-11-11 === * 06:42 [[gitlab:seanhoyland|@seanhoyland]] was approved. === 2025-11-10 === * 00:06 [[gitlab:jaredblumer|@jaredblumer]] was approved. === 2025-11-09 === * 22:36 "heinxiety" was rejected (pending since 2025-08-10T22:33:12.041Z). === 2025-11-07 === * 22:00 [[gitlab:forzagreen|@forzagreen]] was approved. === 2025-11-06 === * 16:57 [[gitlab:rsilvola|@rsilvola]] was approved. === 2025-11-04 === * 21:24 [[gitlab:devdoingdev|@devdoingdev]] was approved. === 2025-11-03 === * 17:48 "joewaleed98" was rejected (pending since 2025-08-04T17:46:12.191Z). === 2025-11-01 === * 18:00 "eliasempresas" was rejected (pending since 2025-08-02T17:58:04.412Z). === 2025-10-31 === * 18:51 [[gitlab:chaoticenby|@chaoticenby]] was approved. * 04:33 "3ch310n" was rejected (pending since 2025-08-01T04:32:21.982Z). === 2025-10-30 === * 10:03 [[gitlab:tausheefhassan|@tausheefhassan]] was approved. === 2025-10-29 === * 14:54 "theap" was rejected (pending since 2025-07-30T14:52:12.066Z). === 2025-10-28 === * 06:06 [[gitlab:tanbiruzzaman|@tanbiruzzaman]] was approved. === 2025-10-27 === * 07:51 [[gitlab:jmoore111|@jmoore111]] was approved. === 2025-10-25 === * 21:09 [[gitlab:valor|@valor]] was approved. * 21:03 [[gitlab:booksmurf|@booksmurf]] was approved. * 02:48 "mystyc1" was rejected (pending since 2025-07-26T02:46:19.373Z). === 2025-10-24 === * 05:12 "aadarshmahesh" was rejected (pending since 2025-07-25T05:09:38.264Z). === 2025-10-22 === * 20:54 [[gitlab:janewanga|@janewanga]] was approved. * 17:27 "abeljeevan" was rejected (pending since 2025-07-23T17:26:46.884Z). * 16:12 "shrimpnaur" was rejected (pending since 2025-07-23T16:10:37.864Z). === 2025-10-21 === * 18:51 "jrmuizel" was rejected (pending since 2025-07-22T18:50:07.315Z). * 09:33 [[gitlab:dpogorzelski|@dpogorzelski]] was approved. === 2025-10-17 === * 13:21 [[gitlab:blegodwin|@blegodwin]] was approved. === 2025-10-16 === * 14:51 [[gitlab:bahago|@bahago]] was approved. * 14:12 "harikrishna0005" was rejected (pending since 2025-07-17T14:10:48.385Z). * 14:09 "gauthammohanraj" was rejected (pending since 2025-07-17T14:08:47.643Z). === 2025-10-15 === * 13:48 [[gitlab:adwivedii|@adwivedii]] was approved. * 13:18 [[gitlab:kimbrenekakande|@kimbrenekakande]] was approved. * 13:03 "childmnajennifer" was rejected (pending since 2025-07-16T13:01:50.236Z). * 05:06 "vssb4214" was rejected (pending since 2025-07-16T05:05:33.985Z). === 2025-10-14 === * 19:39 [[gitlab:afanyulionel|@afanyulionel]] was approved. * 15:33 [[gitlab:sadrettin|@sadrettin]] was approved. * 14:18 [[gitlab:tmwyk|@tmwyk]] was approved. * 08:42 "yasu0796" was rejected (pending since 2025-07-15T08:41:26.453Z). === 2025-10-13 === * 16:09 [[gitlab:atlas0007|@atlas0007]] was approved. === 2025-10-11 === * 17:42 [[gitlab:techwizzie|@techwizzie]] was approved. === 2025-10-10 === * 19:03 [[gitlab:miiswom|@miiswom]] was approved. * 16:06 [[gitlab:ninatakang|@ninatakang]] was approved. === 2025-10-09 === * 15:42 [[gitlab:jaykaneki|@jaykaneki]] was approved. * 14:21 [[gitlab:lebogang|@lebogang]] was approved. * 14:15 [[gitlab:kimondorose|@kimondorose]] was approved. * 13:48 [[gitlab:joyakinyi|@joyakinyi]] was approved. * 13:48 [[gitlab:dikshyashahi|@dikshyashahi]] was approved. * 13:45 [[gitlab:obediobadiah|@obediobadiah]] was approved. * 13:45 [[gitlab:system625|@system625]] was approved. * 13:45 [[gitlab:rolalove|@rolalove]] was approved. * 13:39 [[gitlab:olatundeawo|@olatundeawo]] was approved. * 13:36 [[gitlab:danielchristlight|@danielchristlight]] was approved. * 13:36 [[gitlab:dipanshu1223|@dipanshu1223]] was approved. * 13:36 [[gitlab:aradhya|@aradhya]] was approved. * 09:57 "bognd" was rejected (pending since 2025-07-10T09:55:48.661Z). === 2025-10-08 === * 23:36 [[gitlab:sopzy|@sopzy]] was approved. * 23:03 [[gitlab:oluwatumininu|@oluwatumininu]] was approved. * 19:39 [[gitlab:levon003|@levon003]] was approved. * 15:24 [[gitlab:ritika-bhambri11|@ritika-bhambri11]] was approved. * 13:45 [[gitlab:anbanguyen|@anbanguyen]] was approved. * 13:36 [[gitlab:chumzine|@chumzine]] was approved. * 13:27 [[gitlab:shr0x-ya|@shr0x-ya]] was approved. * 12:45 [[gitlab:nurahwakili|@nurahwakili]] was approved. * 03:42 "nazhiba" was rejected (pending since 2025-07-09T03:40:12.625Z). * 02:12 "mafennel" was rejected (pending since 2025-07-09T02:11:40.598Z). === 2025-10-07 === * 22:54 [[gitlab:olusegunfaj|@olusegunfaj]] was approved. * 21:30 [[gitlab:rona|@rona]] was approved. * 21:09 [[gitlab:sandijigs|@sandijigs]] was approved. * 13:36 "xisbajao" was rejected (pending since 2025-07-08T13:33:35.018Z). * 01:36 "areczek94" was rejected (pending since 2025-07-08T01:35:40.633Z). === 2025-10-06 === * 19:21 "wmcarter2017" was rejected (pending since 2025-07-07T19:21:12.899Z). === 2025-10-05 === * 14:15 "meetmendapara" was rejected (pending since 2025-07-06T14:14:16.726Z). === 2025-10-04 === * 20:51 "nftbaee" was rejected (pending since 2025-07-05T20:50:57.688Z). === 2025-10-03 === * 06:12 [[gitlab:javiermonton|@javiermonton]] was approved. === 2025-10-02 === * 20:15 "talaqalotaibipmp" was rejected (pending since 2025-07-03T20:13:05.164Z). === 2025-10-01 === * 10:54 "bjensen" was rejected (pending since 2025-07-02T10:53:46.574Z). * 02:45 "kowal1984" was rejected (pending since 2025-07-02T02:44:56.946Z). === 2025-09-30 === * 21:21 [[gitlab:kavaljeetsingh|@kavaljeetsingh]] was approved. * 00:24 "adium" was rejected (pending since 2025-07-01T00:23:43.807Z). === 2025-09-28 === * 08:54 [[gitlab:pexerik|@pexerik]] was approved. === 2025-09-27 === * 13:57 [[gitlab:rubahhitamvukova|@rubahhitamvukova]] was approved. === 2025-09-26 === * 16:57 "algorithmic" was rejected (pending since 2025-06-27T16:56:17.480Z). * 13:54 [[gitlab:shadabgdg|@shadabgdg]] was approved. * 13:12 [[gitlab:spushpit|@spushpit]] was approved. === 2025-09-20 === * 14:06 "bwiki" was rejected (pending since 2025-06-21T13:59:14.749Z). === 2025-09-16 === * 05:39 [[gitlab:deepchirp|@deepchirp]] was approved. === 2025-09-15 === * 22:00 [[gitlab:noisk8|@noisk8]] was approved. * 11:03 "ahonc" was rejected (pending since 2025-06-16T11:00:54.843Z). === 2025-09-13 === * 18:24 "a-ssh22" was rejected (pending since 2025-06-14T18:23:33.937Z). * 12:36 [[gitlab:rajashreetalukdar|@rajashreetalukdar]] was approved. * 00:45 [[gitlab:sumitsurai|@sumitsurai]] was approved. === 2025-09-12 === * 17:12 [[gitlab:suyash23|@suyash23]] was approved. * 00:46 "remotetravel" was rejected (pending since 2025-06-13T00:44:08.171Z). === 2025-09-10 === * 21:09 "jancborchardt" was rejected (pending since 2025-06-11T21:06:30.759Z). === 2025-09-09 === * 17:03 [[gitlab:vwf|@vwf]] was approved. * 06:36 [[gitlab:cactusisme|@cactusisme]] was approved. === 2025-09-08 === * 18:09 "birushandegeya" was rejected (pending since 2025-06-09T18:08:00.087Z). * 16:27 "ngarnsworthy" was rejected (pending since 2025-06-09T16:24:37.213Z). * 12:33 "zolgoyo" was rejected (pending since 2025-06-09T12:31:34.199Z). === 2025-09-06 === * 23:09 [[gitlab:jaishsingh913|@jaishsingh913]] was approved. === 2025-09-05 === * 21:45 [[gitlab:sakshi2|@sakshi2]] was approved. * 20:42 "abdukhaliq1" was rejected (pending since 2025-06-06T20:40:42.023Z). * 14:27 "beubsamy" was rejected (pending since 2025-06-06T14:27:06.781Z). === 2025-09-04 === * 23:27 "sdhehua" was rejected (pending since 2025-06-05T23:24:45.777Z). * 19:00 [[gitlab:perry|@perry]] was approved. * 11:24 "saintwolf" was rejected (pending since 2025-06-05T11:21:20.176Z). === 2025-09-02 === * 05:48 [[gitlab:aliu|@aliu]] was approved. === 2025-08-29 === * 13:30 "kksurendran066" was rejected (pending since 2025-05-30T13:27:48.755Z). === 2025-08-28 === * 22:18 "tauraamuix" was rejected (pending since 2025-05-29T22:16:08.228Z). === 2025-08-26 === * 19:03 [[gitlab:dikkulah|@dikkulah]] was approved. === 2025-08-22 === * 23:51 [[gitlab:khoroshun_mike|@khoroshun_mike]] was approved. === 2025-08-21 === * 07:39 [[gitlab:yuka|@yuka]] was approved. === 2025-08-19 === * 07:48 [[gitlab:zhaofjx|@zhaofjx]] was approved. === 2025-08-17 === * 14:27 "madhan13k" was rejected (pending since 2025-05-18T14:26:08.973Z). === 2025-08-15 === * 10:15 "mohammed_abukhadra" was rejected (pending since 2025-05-16T10:14:48.403Z). === 2025-08-11 === * 11:48 "hmmyesbro" was rejected (pending since 2025-05-12T11:45:24.350Z). === 2025-08-10 === * 13:15 [[gitlab:dactyl|@dactyl]] was approved. === 2025-08-09 === * 04:39 "xxxx100000" was rejected (pending since 2025-05-10T04:37:44.949Z). === 2025-08-08 === * 14:33 [[gitlab:josefanthony|@josefanthony]] was approved. === 2025-08-07 === * 23:42 [[gitlab:robins7|@robins7]] was approved. * 21:42 [[gitlab:pols12|@pols12]] was approved. * 17:15 "sbronson" was rejected (pending since 2025-05-08T17:15:08.834Z). * 14:57 [[gitlab:alvindulle|@alvindulle]] was approved. * 14:45 [[gitlab:xentos|@xentos]] was approved. * 06:27 "jamesboste" was rejected (pending since 2025-05-08T06:25:14.793Z). * 03:57 "ysun" was rejected (pending since 2025-05-08T03:55:07.348Z). === 2025-08-06 === * 21:51 "pols12" was rejected (pending since 2025-05-07T21:49:13.598Z). * 01:51 "okeamah" was rejected (pending since 2025-05-07T01:48:50.114Z). === 2025-08-05 === * 09:15 "mobashir-2013" was rejected (pending since 2025-05-06T09:14:24.069Z). === 2025-08-01 === * 08:00 "douginamug" was rejected (pending since 2025-05-02T07:57:38.317Z). === 2025-07-31 === * 02:30 [[gitlab:ads|@ads]] was approved. === 2025-07-27 === * 13:15 "mrico2703" was rejected (pending since 2025-04-27T13:13:12.346Z). * 10:17 [[gitlab:josephfrancis12|@josephfrancis12]] was approved. * 10:17 [[gitlab:fuzzew|@fuzzew]] was approved. * 05:57 [[gitlab:biscuitbobby|@biscuitbobby]] was approved. * 05:48 [[gitlab:ecoholic|@ecoholic]] was approved. === 2025-07-26 === * 11:48 [[gitlab:chimnayyyy|@chimnayyyy]] was approved. * 11:48 [[gitlab:alwinalbert|@alwinalbert]] was approved. * 11:48 [[gitlab:hridyakk|@hridyakk]] was approved. * 11:45 [[gitlab:gaurigupta21|@gaurigupta21]] was approved. * 11:45 [[gitlab:binetaa|@binetaa]] was approved. * 10:21 [[gitlab:jyothikat22|@jyothikat22]] was approved. * 10:21 [[gitlab:zobotrombie|@zobotrombie]] was approved. * 10:21 [[gitlab:flykrth|@flykrth]] was approved. * 10:21 [[gitlab:mehrinshamim|@mehrinshamim]] was approved. * 10:21 [[gitlab:aadhi13|@aadhi13]] was approved. * 10:21 [[gitlab:malavikam05|@malavikam05]] was approved. * 10:18 [[gitlab:nf609|@nf609]] was approved. * 05:48 [[gitlab:nazalnihad|@nazalnihad]] was approved. * 05:48 [[gitlab:naveen28204280|@naveen28204280]] was approved. === 2025-07-25 === * 09:49 [[gitlab:kasyap9|@kasyap9]] was approved. * 09:30 [[gitlab:swayamagrahari|@swayamagrahari]] was approved. === 2025-07-24 === * 19:36 [[gitlab:madutgn|@madutgn]] was approved. === 2025-07-23 === * 20:09 [[gitlab:somerandomdeveloper|@somerandomdeveloper]] was approved. === 2025-07-22 === * 00:15 [[gitlab:iagoqnsi|@iagoqnsi]] was approved. === 2025-07-21 === * 17:30 [[gitlab:asadiqui|@asadiqui]] was approved. * 16:39 [[gitlab:tryvix1509|@tryvix1509]] was approved. * 04:27 [[gitlab:damian|@damian]] was approved. === 2025-07-20 === * 09:42 "mike-khoroshun" was rejected (pending since 2025-04-20T09:42:22.732Z). === 2025-07-17 === * 17:57 [[gitlab:haroldkrabs|@haroldkrabs]] was approved. * 13:45 [[gitlab:envlh|@envlh]] was approved. === 2025-07-14 === * 10:24 [[gitlab:missguru|@missguru]] was approved. * 00:57 "clarfonthey" was rejected (pending since 2025-04-14T00:56:32.626Z). === 2025-07-13 === * 01:01 [[gitlab:l235|@l235]] was approved. === 2025-07-11 === * 03:06 "rodavlas" was rejected (pending since 2025-04-11T03:05:45.590Z). === 2025-07-06 === * 00:09 "lakasa" was rejected (pending since 2025-04-06T00:06:28.469Z). === 2025-07-05 === * 21:54 "ctrlzvi" was rejected (pending since 2025-04-05T21:54:12.542Z). * 14:30 "aminualiyu" was rejected (pending since 2025-04-05T14:27:22.617Z). === 2025-07-04 === * 03:15 [[gitlab:galstar|@galstar]] was approved. === 2025-07-02 === * 11:27 "vicolas11" was rejected (pending since 2025-04-02T11:25:12.682Z). === 2025-06-29 === * 23:12 "naomi723" was rejected (pending since 2025-03-30T23:09:24.630Z). === 2025-06-28 === * 16:21 "mudeh2372" was rejected (pending since 2025-03-29T16:18:27.057Z). === 2025-06-27 === * 23:18 "rony143" was rejected (pending since 2025-03-28T23:16:13.671Z). * 22:21 [[gitlab:rluts|@rluts]] was approved. === 2025-06-26 === * 13:54 "creativegurus" was rejected (pending since 2025-03-27T13:52:41.706Z). === 2025-06-24 === * 17:42 [[gitlab:devjadiya|@devjadiya]] was approved. * 14:00 "dominic-r" was rejected (pending since 2025-03-25T14:00:07.307Z). === 2025-06-21 === * 00:48 [[gitlab:vriaa|@vriaa]] was approved. === 2025-06-18 === * 15:21 "ayushkhati1" was rejected (pending since 2025-03-19T15:18:50.062Z). === 2025-06-17 === * 20:45 "chiomavero" was rejected (pending since 2025-03-18T20:44:13.967Z). * 00:27 [[gitlab:eggroll97|@eggroll97]] was approved. === 2025-06-14 === * 20:57 "volvox" was rejected (pending since 2025-03-15T20:56:34.018Z). === 2025-06-13 === * 16:09 [[gitlab:supergrey|@supergrey]] was approved. * 11:03 "chqaz" was rejected (pending since 2025-03-14T11:01:09.600Z). * 10:24 [[gitlab:slong-wmf|@slong-wmf]] was approved. * 10:15 "hearvox" was rejected (pending since 2025-03-14T10:13:13.112Z). === 2025-06-12 === * 15:18 "jlam" was rejected (pending since 2025-03-13T15:17:54.099Z). === 2025-06-09 === * 20:48 "dipanjansengupta" was rejected (pending since 2025-03-10T20:48:03.545Z). * 19:27 [[gitlab:reggycelly|@reggycelly]] was approved. * 14:51 "arendpieter" was rejected (pending since 2025-03-10T14:51:01.445Z). * 13:21 [[gitlab:greenreaper|@greenreaper]] was approved. * 09:33 [[gitlab:mmta|@mmta]] was approved. * 08:03 "a-ssh22" was rejected (pending since 2025-03-10T08:03:08.111Z). === 2025-06-08 === * 21:06 "mm-episodenlistedlvaupdater" was rejected (pending since 2025-03-09T21:04:06.323Z). === 2025-06-06 === * 11:06 [[gitlab:olea|@olea]] was approved. === 2025-06-05 === * 20:33 [[gitlab:encodedwp|@encodedwp]] was approved. * 15:00 [[gitlab:toluayo|@toluayo]] was approved. * 13:51 [[gitlab:arnold_lup|@arnold_lup]] was approved. * 11:54 "sdhehua" was rejected (pending since 2025-03-06T11:51:48.241Z). === 2025-06-03 === * 21:27 [[gitlab:wewakey|@wewakey]] was approved. * 12:36 "hunsimon2" was rejected (pending since 2025-03-04T12:34:56.520Z). * 11:54 "hunsimon" was rejected (pending since 2025-03-04T11:53:54.652Z). === 2025-06-02 === * 12:01 [[gitlab:jaimedes|@jaimedes]] was approved. === 2025-05-30 === * 18:00 "sathvik9105" was rejected (pending since 2025-02-28T17:59:42.867Z). * 11:21 [[gitlab:tonythomas01|@tonythomas01]] was approved. * 10:06 [[gitlab:gpsleo|@gpsleo]] was approved. === 2025-05-29 === * 22:12 [[gitlab:codynguyen1116|@codynguyen1116]] was approved. === 2025-05-28 === * 02:57 [[gitlab:saper|@saper]] was approved. === 2025-05-27 === * 21:06 [[gitlab:mohammed_qays|@mohammed_qays]] was approved. * 15:33 "satanluimm" was rejected (pending since 2025-02-25T15:32:48.101Z). === 2025-05-26 === * 23:57 "seyedali220" was rejected (pending since 2025-02-24T23:56:17.621Z). === 2025-05-21 === * 11:12 [[gitlab:guilherme|@guilherme]] was approved. === 2025-05-19 === * 13:24 [[gitlab:emojiwiki|@emojiwiki]] was approved. === 2025-05-18 === * 00:00 "xidme" was rejected (pending since 2025-02-15T23:58:56.796Z). === 2025-05-17 === * 02:39 "kdh8219" was rejected (pending since 2025-02-15T02:36:32.237Z). === 2025-05-16 === * 15:09 [[gitlab:maxbinderwmf|@maxbinderwmf]] was approved. === 2025-05-15 === * 04:30 "inspectorzer0" was rejected (pending since 2025-02-13T04:27:33.179Z). === 2025-05-14 === * 17:42 [[gitlab:llugo|@llugo]] was approved. === 2025-05-13 === * 20:18 "mmta" was rejected (pending since 2025-02-11T20:17:23.407Z). === 2025-05-11 === * 20:51 "jad" was rejected (pending since 2025-02-09T20:49:07.333Z). * 17:54 "nishchalsundan" was rejected (pending since 2025-02-09T17:52:25.761Z). * 16:39 "mohammed_abukhadra" was rejected (pending since 2025-02-09T16:39:03.730Z). === 2025-05-09 === * 09:12 [[gitlab:sirchanmp|@sirchanmp]] was approved. === 2025-05-08 === * 08:18 [[gitlab:mengeditch|@mengeditch]] was approved. === 2025-05-07 === * 03:45 "xluffy" was rejected (pending since 2025-02-05T03:45:14.181Z). === 2025-05-06 === * 16:54 "punhaniabhishek" was rejected (pending since 2025-02-04T16:53:50.758Z). * 09:36 [[gitlab:bmartinezcalvo|@bmartinezcalvo]] was approved. === 2025-05-02 === * 12:24 [[gitlab:tohaomg|@tohaomg]] was approved. * 11:48 [[gitlab:mavrikant|@mavrikant]] was approved. * 11:45 [[gitlab:daanvr|@daanvr]] was approved. === 2025-05-01 === * 09:09 "mjoerg" was rejected (pending since 2025-01-30T09:09:04.204Z). === 2025-04-30 === * 23:06 "sanskardubey" was rejected (pending since 2025-01-29T23:03:25.489Z). === 2025-04-29 === * 16:00 "geyslein" was rejected (pending since 2025-01-28T16:00:01.510Z). === 2025-04-26 === * 09:30 "anjali9027" was rejected (pending since 2025-01-25T09:28:07.064Z). === 2025-04-25 === * 18:00 "salahhazaa" was rejected (pending since 2025-01-24T17:58:30.030Z). * 15:15 [[gitlab:yiming|@yiming]] was approved. * 02:06 "mrchanmp" was rejected (pending since 2025-01-24T02:03:58.308Z). === 2025-04-23 === * 17:03 "rj2904" was rejected (pending since 2025-01-22T17:03:11.207Z). * 14:21 "nischay33" was rejected (pending since 2025-01-22T14:19:21.081Z). === 2025-04-22 === * 19:27 "dj80" was rejected (pending since 2025-01-21T19:25:28.498Z). * 14:30 [[gitlab:kaimamin|@kaimamin]] was approved. * 09:57 "debo" was rejected (pending since 2025-01-21T09:54:47.955Z). === 2025-04-21 === * 12:24 "unshell" was rejected (pending since 2025-01-20T12:21:59.686Z). === 2025-04-18 === * 15:06 [[gitlab:spartanarbinger|@spartanarbinger]] was approved. === 2025-04-16 === * 03:09 "dewey" was rejected (pending since 2025-01-15T03:06:17.488Z). === 2025-04-15 === * 19:45 "emdadul" was rejected (pending since 2025-01-14T19:42:29.285Z). === 2025-04-14 === * 06:45 [[gitlab:bcampbell804|@bcampbell804]] was approved. === 2025-04-11 === * 06:27 [[gitlab:jvanderhoop|@jvanderhoop]] was approved. === 2025-04-10 === * 04:12 "bhai420" was rejected (pending since 2025-01-09T04:10:29.430Z). === 2025-04-09 === * 05:03 "austinvarshney" was rejected (pending since 2025-01-08T05:02:34.175Z). === 2025-04-06 === * 15:36 [[gitlab:elph|@elph]] was approved. === 2025-04-02 === * 10:33 [[gitlab:ozge|@ozge]] was approved. === 2025-03-31 === * 20:15 "demandkey" was rejected (pending since 2024-12-30T20:14:23.096Z). * 15:18 [[gitlab:danyya|@danyya]] was approved. === 2025-03-28 === * 15:54 [[gitlab:rutsavi09|@rutsavi09]] was approved. * 15:54 [[gitlab:ilanen1|@ilanen1]] was approved. === 2025-03-25 === * 19:27 [[gitlab:irfo|@irfo]] was approved. * 11:54 [[gitlab:kmontalva-wmf|@kmontalva-wmf]] was approved. * 04:33 [[gitlab:paul26|@paul26]] was approved. * 04:18 "as1100k" was rejected (pending since 2024-12-24T04:18:06.813Z). === 2025-03-24 === * 11:33 "amzadkhankk" was rejected (pending since 2024-12-23T11:33:14.176Z). === 2025-03-23 === * 12:24 "wolfdo" was rejected (pending since 2024-12-22T12:23:35.056Z). === 2025-03-22 === * 09:45 [[gitlab:fjmustak|@fjmustak]] was approved. === 2025-03-20 === * 18:42 "sathishkokila" was rejected (pending since 2024-12-19T18:39:35.161Z). * 17:03 [[gitlab:alien4444|@alien4444]] was approved. * 15:27 [[gitlab:davidcoronel|@davidcoronel]] was approved. === 2025-03-19 === * 22:57 [[gitlab:r1f4t|@r1f4t]] was approved. * 19:03 "daniel24ps" was rejected (pending since 2024-12-18T19:00:21.249Z). * 14:18 [[gitlab:beepbooppenguin|@beepbooppenguin]] was approved. === 2025-03-18 === * 17:48 "rahulkundu1209" was rejected (pending since 2024-12-17T17:46:41.936Z). * 08:15 "kirtisikka972" was rejected (pending since 2024-12-17T08:13:25.487Z). === 2025-03-15 === * 13:30 "tulspal_sidhu" was rejected (pending since 2024-12-14T13:29:10.606Z). * 01:39 "peacedeadc" was rejected (pending since 2024-12-14T01:37:36.579Z). === 2025-03-14 === * 03:51 [[gitlab:chuckthebuck|@chuckthebuck]] was approved. * 02:33 "yxngtrtxll" was rejected (pending since 2024-12-13T02:31:51.658Z). === 2025-03-13 === * 14:36 [[gitlab:iccander|@iccander]] was approved. === 2025-03-12 === * 23:21 "jokerchic36" was rejected (pending since 2024-12-11T23:21:00.670Z). * 15:30 [[gitlab:naomi|@naomi]] was approved. * 15:27 [[gitlab:cobi|@cobi]] was approved. === 2025-03-11 === * 12:42 "mohitvermaxx" was rejected (pending since 2024-12-10T12:40:56.967Z). === 2025-03-10 === * 16:51 [[gitlab:nanona15dobato|@nanona15dobato]] was approved. === 2025-03-09 === * 22:39 [[gitlab:jonkolbert|@jonkolbert]] was approved. * 20:45 [[gitlab:urbanecmtest2|@urbanecmtest2]] was approved. === 2025-03-07 === * 16:54 [[gitlab:hswan|@hswan]] was approved. * 14:42 [[gitlab:atitkov|@atitkov]] was approved. * 00:42 [[gitlab:infrastruktur|@infrastruktur]] was approved. === 2025-03-06 === * 17:21 "johnmann" was rejected (pending since 2024-12-05T17:19:24.995Z). === 2025-03-05 === * 07:33 [[gitlab:monx9494|@monx9494]] was approved. === 2025-03-02 === * 21:21 "paul26" was rejected (pending since 2024-12-01T21:20:19.681Z). === 2025-03-01 === * 19:15 [[gitlab:izno|@izno]] was approved. * 12:45 [[gitlab:nyerho|@nyerho]] was approved. === 2025-02-28 === * 18:27 [[gitlab:chuckonwumelu|@chuckonwumelu]] was approved. * 13:09 "ashwinpraveengo" was rejected (pending since 2024-11-29T13:07:47.240Z). * 00:18 "eduardoaugusto" was rejected (pending since 2024-11-29T00:17:43.372Z). === 2025-02-27 === * 20:39 "volkanurl" was rejected (pending since 2024-11-28T20:37:18.101Z). === 2025-02-24 === * 21:15 [[gitlab:feeglgeef|@feeglgeef]] was approved. * 20:18 [[gitlab:piaanalysis2|@piaanalysis2]] was approved. * 19:06 [[gitlab:dhardy|@dhardy]] was approved. === 2025-02-22 === * 19:27 [[gitlab:owuh|@owuh]] was approved. === 2025-02-19 === * 16:06 [[gitlab:artemkloko|@artemkloko]] was approved. * 13:03 [[gitlab:jgafnea|@jgafnea]] was approved. === 2025-02-17 === * 16:33 [[gitlab:asmartkitten|@asmartkitten]] was approved. === 2025-02-16 === * 19:12 "gaurigupta21" was rejected (pending since 2024-11-17T19:11:07.416Z). === 2025-02-15 === * 01:18 [[gitlab:mediawiki-quickstart-ci|@mediawiki-quickstart-ci]] was approved. === 2025-02-14 === * 15:21 "nathanbnm" was rejected (pending since 2024-11-15T15:18:19.632Z). === 2025-02-13 === * 16:45 [[gitlab:priyanshuchahal|@priyanshuchahal]] was approved. * 16:42 [[gitlab:ajhalili2006|@ajhalili2006]] was approved. === 2025-02-12 === * 23:21 "monkeypatch999" was rejected (pending since 2024-11-13T23:20:38.398Z). * 06:36 [[gitlab:jainlakshita28|@jainlakshita28]] was approved. === 2025-02-11 === * 19:27 [[gitlab:matthewsm2|@matthewsm2]] was approved. === 2025-02-09 === * 16:15 "mohammed_abukhadra" was rejected (pending since 2024-11-10T16:15:18.361Z). === 2025-02-07 === * 21:33 "brennan" was rejected (pending since 2024-11-08T21:31:07.351Z). === 2025-02-06 === * 08:24 "mmta" was rejected (pending since 2024-11-07T08:22:36.724Z). * 06:21 [[gitlab:bunnypranav|@bunnypranav]] was approved. === 2025-02-05 === * 22:39 "chrissteinchen" was rejected (pending since 2024-11-06T22:38:16.673Z). === 2025-02-03 === * 07:45 "edriiic" was rejected (pending since 2024-11-04T07:44:46.849Z). * 01:12 "geppy" was rejected (pending since 2024-11-04T01:10:48.710Z). === 2025-02-02 === * 13:18 "funa-enpitu" was rejected (pending since 2024-11-03T13:15:46.065Z). === 2025-01-31 === * 23:42 "nfontes" was rejected (pending since 2024-11-01T23:39:41.755Z). * 22:51 "sbronson" was rejected (pending since 2024-11-01T22:50:31.871Z). * 00:42 [[gitlab:farid|@farid]] was approved. === 2025-01-27 === * 08:15 [[gitlab:eliza189|@eliza189]] was approved. === 2025-01-25 === * 09:51 [[gitlab:pamputt|@pamputt]] was approved. === 2025-01-23 === * 14:30 [[gitlab:lubianat|@lubianat]] was approved. * 11:45 [[gitlab:bootsa|@bootsa]] was approved. === 2025-01-21 === * 05:09 "niko" was rejected (pending since 2024-07-21T16:10:01.377Z). * 05:09 "thawizkid369777" was rejected (pending since 2024-07-18T17:42:44.493Z). * 05:09 "sarthaksingh2" was rejected (pending since 2024-07-10T11:31:30.470Z). * 05:09 "shriyakt" was rejected (pending since 2024-07-06T04:54:10.248Z). * 05:09 "akshaya" was rejected (pending since 2024-07-06T04:04:51.488Z). * 05:09 "alaka03aj" was rejected (pending since 2024-07-05T18:01:54.876Z). * 05:09 "sulochanaviji-5049" was rejected (pending since 2024-07-01T05:58:00.427Z). * 05:09 "nayanjnath" was rejected (pending since 2024-07-01T02:51:57.405Z). * 05:09 "sd44" was rejected (pending since 2024-06-30T04:28:51.436Z). * 05:09 "metavalent" was rejected (pending since 2024-06-29T01:37:14.210Z). * 05:09 "wicloudx" was rejected (pending since 2024-06-28T11:51:23.335Z). * 05:09 "debo" was rejected (pending since 2024-06-28T01:44:59.845Z). * 05:09 "bwiki" was rejected (pending since 2024-06-23T14:15:38.032Z). * 05:09 "toprak" was rejected (pending since 2024-06-23T11:35:50.819Z). * 05:09 "iristeller" was rejected (pending since 2024-06-14T20:53:48.959Z). * 05:09 "jcolvin" was rejected (pending since 2024-06-12T17:29:01.238Z). * 05:09 "kalyan" was rejected (pending since 2024-06-07T07:52:46.993Z). * 05:09 "bluecrystal" was rejected (pending since 2024-06-06T19:16:20.107Z). * 05:09 "iftttrohit" was rejected (pending since 2024-06-04T12:08:50.818Z). * 05:09 "pogpotato" was rejected (pending since 2024-06-03T17:58:21.684Z). * 05:09 "cptlausebaer" was rejected (pending since 2024-05-31T18:53:27.692Z). * 05:09 "hdevine825" was rejected (pending since 2024-05-31T17:04:18.279Z). * 05:09 "anaghaa18" was rejected (pending since 2024-05-25T19:14:31.803Z). * 05:09 "atharvanair04" was rejected (pending since 2024-05-25T14:24:52.825Z). * 05:09 "anasvemmully" was rejected (pending since 2024-05-25T06:10:27.261Z). * 05:09 "abhinavmohandas" was rejected (pending since 2024-05-25T06:05:24.825Z). * 05:09 "kksurendran06" was rejected (pending since 2024-05-25T06:04:38.082Z). * 05:09 "albertmarshall8896" was rejected (pending since 2024-05-23T09:32:05.462Z). * 05:09 "akellison" was rejected (pending since 2024-05-17T02:07:24.229Z). * 05:09 "mainowill" was rejected (pending since 2024-04-16T23:30:33.881Z). * 05:09 "bzhqc" was rejected (pending since 2024-04-16T19:50:38.676Z). * 05:09 "safan41" was rejected (pending since 2024-04-16T03:34:48.942Z). * 05:09 "mgagat" was rejected (pending since 2024-04-16T03:21:51.764Z). * 05:09 "okeamah" was rejected (pending since 2024-04-16T02:49:00.143Z). * 05:09 "xuhao61" was rejected (pending since 2024-04-15T23:45:09.083Z). * 04:47 "cybel" was rejected (pending since 2024-04-15T06:46:35.791Z). === 2025-01-20 === * 14:33 [[gitlab:your1|@your1]] was approved. === 2025-01-18 === * 10:09 [[gitlab:galrach600|@galrach600]] was approved. * 02:51 [[gitlab:blankeclair|@blankeclair]] was approved. === 2025-01-17 === * 13:57 [[gitlab:dsantamaria|@dsantamaria]] was approved. === 2025-01-15 === * 17:12 [[gitlab:smartse|@smartse]] was approved. === 2025-01-14 === * 17:03 [[gitlab:naorleizer|@naorleizer]] was approved. === 2025-01-13 === * 02:45 [[gitlab:wolf20482|@wolf20482]] was approved. === 2025-01-12 === * 17:45 [[gitlab:tamzin|@tamzin]] was approved. === 2025-01-11 === * 15:24 [[gitlab:bargioni|@bargioni]] was approved. * 14:30 [[gitlab:salelya|@salelya]] was approved. * 10:15 [[gitlab:malakatshy|@malakatshy]] was approved. * 05:21 [[gitlab:newmcpee|@newmcpee]] was approved. === 2025-01-09 === * 15:30 [[gitlab:gkyziridis|@gkyziridis]] was approved. === 2025-01-08 === * 16:21 [[gitlab:ukrface|@ukrface]] was approved. === 2024-12-28 === * 03:27 [[gitlab:twonum|@twonum]] was approved. === 2024-12-25 === * 06:09 [[gitlab:harsv567|@harsv567]] was approved. === 2024-12-21 === * 11:24 [[gitlab:amutha2002|@amutha2002]] was approved. === 2024-12-20 === * 19:51 [[gitlab:hridyeshgupta|@hridyeshgupta]] was approved. * 10:00 [[gitlab:ro-shines|@ro-shines]] was approved. * 08:09 [[gitlab:kesharwaniarpita|@kesharwaniarpita]] was approved. === 2024-12-18 === * 14:45 [[gitlab:soylacarli|@soylacarli]] was approved. === 2024-12-16 === * 20:33 [[gitlab:aleyasiddika1|@aleyasiddika1]] was approved. === 2024-12-15 === * 07:33 [[gitlab:abhishek02bhardwaj|@abhishek02bhardwaj]] was approved. === 2024-12-13 === * 13:18 [[gitlab:ashmitabathre204|@ashmitabathre204]] was approved. === 2024-12-10 === * 06:39 [[gitlab:ginaan|@ginaan]] was approved. === 2024-12-09 === * 05:45 [[gitlab:kallinavya|@kallinavya]] was approved. * 00:54 [[gitlab:viserion-7|@viserion-7]] was approved. === 2024-12-08 === * 17:27 [[gitlab:wargo|@wargo]] was approved. === 2024-12-05 === * 11:15 [[gitlab:ranjithraj|@ranjithraj]] was approved. === 2024-12-02 === * 21:21 [[gitlab:a930913|@a930913]] was approved. === 2024-12-01 === * 02:39 [[gitlab:kingchristlike1|@kingchristlike1]] was approved. === 2024-11-21 === * 13:45 [[gitlab:sascha|@sascha]] was approved. === 2024-11-19 === * 16:36 [[gitlab:jly|@jly]] was approved. === 2024-11-15 === * 02:54 [[gitlab:danielyepezgarces|@danielyepezgarces]] was approved. === 2024-11-14 === * 14:15 [[gitlab:stimoroll|@stimoroll]] was approved. === 2024-11-09 === * 17:15 [[gitlab:f4udeveloper|@f4udeveloper]] was approved. === 2024-11-07 === * 19:15 [[gitlab:zulf|@zulf]] was approved. * 05:33 [[gitlab:hassanamin|@hassanamin]] was approved. === 2024-11-06 === * 19:39 [[gitlab:daniuu|@daniuu]] was approved. * 00:18 [[gitlab:rlopez-wmf|@rlopez-wmf]] was approved. === 2024-10-09 === * 14:45 [[gitlab:jtweed|@jtweed]] was approved. * 10:24 [[gitlab:ifrahkh|@ifrahkh]] was approved. * 09:06 [[gitlab:wikibayer|@wikibayer]] was approved. === 2024-10-06 === * 10:27 [[gitlab:keerthan16|@keerthan16]] was approved. === 2024-10-04 === * 07:45 [[gitlab:hakimi97|@hakimi97]] was approved. === 2024-09-30 === * 07:39 [[gitlab:ninjastrikers|@ninjastrikers]] was approved. === 2024-09-28 === * 17:30 [[gitlab:webrunner95|@webrunner95]] was approved. === 2024-09-18 === * 21:39 [[gitlab:elliottetzkorn|@elliottetzkorn]] was approved. === 2024-09-14 === * 22:06 [[gitlab:humptydumpty|@humptydumpty]] was approved. === 2024-09-06 === * 08:48 [[gitlab:mickabarber|@mickabarber]] was approved. === 2024-08-27 === * 17:36 [[gitlab:edgars|@edgars]] was approved. === 2024-08-22 === * 09:18 [[gitlab:antonkokhwmde|@antonkokhwmde]] was approved. === 2024-08-14 === * 19:21 [[gitlab:jfk|@jfk]] was approved. === 2024-08-13 === * 17:57 [[gitlab:daxserver|@daxserver]] was approved. === 2024-08-11 === * 09:57 [[gitlab:pauliesnug|@pauliesnug]] was approved. === 2024-08-10 === * 08:42 [[gitlab:ashig|@ashig]] was approved. === 2024-08-09 === * 14:09 [[gitlab:masssly|@masssly]] was approved. === 2024-08-05 === * 22:15 [[gitlab:mrtortue|@mrtortue]] was approved. === 2024-08-02 === * 16:21 [[gitlab:dsantini|@dsantini]] was approved. === 2024-07-31 === * 11:54 [[gitlab:cptviraj|@cptviraj]] was approved. === 2024-07-30 === * 19:09 [[gitlab:iniquity|@iniquity]] was approved. * 10:00 [[gitlab:collins|@collins]] was approved. === 2024-07-27 === * 15:57 [[gitlab:songnguxyz|@songnguxyz]] was approved. === 2024-07-25 === * 12:36 [[gitlab:mszabo|@mszabo]] was approved. * 09:21 [[gitlab:agarwalmahima|@agarwalmahima]] was approved. === 2024-07-24 === * 08:05 [[gitlab:dragoniez|@dragoniez]] was approved. === 2024-07-23 === * 06:54 [[gitlab:mirji|@mirji]] was approved. === 2024-07-16 === * 10:00 [[gitlab:lakejason0|@lakejason0]] was approved. === 2024-07-12 === * 11:33 [[gitlab:cn|@cn]] was approved. * 08:12 [[gitlab:unchampignon|@unchampignon]] was approved. === 2024-07-07 === * 17:12 [[gitlab:agamyasamuel|@agamyasamuel]] was approved. * 05:24 [[gitlab:kuldeepburjbhalaike|@kuldeepburjbhalaike]] was approved. === 2024-07-06 === * 11:18 [[gitlab:dibya|@dibya]] was approved. * 04:54 [[gitlab:sarthakparashar|@sarthakparashar]] was approved. === 2024-07-05 === * 18:15 [[gitlab:vanshikarathi|@vanshikarathi]] was approved. === 2024-07-02 === * 19:00 [[gitlab:ebrahim|@ebrahim]] was approved. === 2024-07-01 === * 20:12 [[gitlab:rockingpenny4|@rockingpenny4]] was approved. * 18:15 [[gitlab:balajijagadesh|@balajijagadesh]] was approved. === 2024-06-30 === * 18:24 [[gitlab:hrideshmg|@hrideshmg]] was approved. * 07:18 [[gitlab:chanakyakumardas|@chanakyakumardas]] was approved. * 06:30 [[gitlab:rihaan180|@rihaan180]] was approved. === 2024-06-27 === * 17:36 [[gitlab:driedmueller|@driedmueller]] was approved. === 2024-06-19 === * 12:57 [[gitlab:audreypenven|@audreypenven]] was approved. === 2024-06-16 === * 01:18 [[gitlab:roysmith|@roysmith]] was approved. === 2024-06-08 === * 02:45 [[gitlab:jleedev|@jleedev]] was approved. === 2024-06-03 === * 13:57 [[gitlab:afeder|@afeder]] was approved. === 2024-06-01 === * 10:54 [[gitlab:florianschmitt|@florianschmitt]] was approved. === 2024-05-30 === * 16:42 [[gitlab:krlsca|@krlsca]] was approved. === 2024-05-28 === * 11:24 [[gitlab:rickijay|@rickijay]] was approved. === 2024-05-26 === * 11:18 [[gitlab:ranjithsiji|@ranjithsiji]] was approved. === 2024-05-25 === * 07:24 [[gitlab:jony|@jony]] was approved. === 2024-05-23 === * 08:45 [[gitlab:lepticed7|@lepticed7]] was approved. === 2024-05-22 === * 20:42 [[gitlab:echecs|@echecs]] was approved. === 2024-05-21 === * 13:33 [[gitlab:mbs|@mbs]] was approved. === 2024-05-19 === * 18:06 [[gitlab:ionenlaser|@ionenlaser]] was approved. === 2024-05-18 === * 23:36 [[gitlab:mdaniels5757|@mdaniels5757]] was approved. === 2024-05-17 === * 08:54 [[gitlab:grapedog|@grapedog]] was approved. === 2024-05-08 === * 19:42 [[gitlab:kelhurd|@kelhurd]] was approved. * 19:06 [[gitlab:khurd|@khurd]] was approved. === 2024-05-06 === * 19:48 [[gitlab:j3j5|@j3j5]] was approved. * 12:06 [[gitlab:tk-999|@tk-999]] was approved. === 2024-05-05 === * 22:09 [[gitlab:pppery|@pppery]] was approved. * 20:33 [[gitlab:sakretsu|@sakretsu]] was approved. * 12:12 [[gitlab:waterquark|@waterquark]] was approved. === 2024-05-04 === * 09:03 [[gitlab:multichill|@multichill]] was approved. * 07:42 [[gitlab:abaris|@abaris]] was approved. === 2024-05-03 === * 14:57 [[gitlab:maurusian|@maurusian]] was approved. === 2024-04-24 === * 05:48 [[gitlab:wolfinux|@wolfinux]] was approved. === 2024-04-23 === * 15:48 [[gitlab:dreamrimmer|@dreamrimmer]] was approved. === 2024-04-21 === * 06:51 [[gitlab:alon|@alon]] was approved. === 2024-04-17 === * 23:33 [[gitlab:derenrich|@derenrich]] was approved. === 2024-04-16 === * 17:18 [[gitlab:valcio|@valcio]] was approved. === 2024-04-14 === * 16:51 [[gitlab:wikilucas00|@wikilucas00]] was approved. === 2024-04-06 === * 12:48 [[gitlab:theprotonade|@theprotonade]] was approved. === 2024-04-02 === * 07:30 [[gitlab:bohuizhang|@bohuizhang]] was approved. === 2024-03-30 === * 13:36 [[gitlab:lpintscher|@lpintscher]] was approved. === 2024-03-26 === * 17:09 [[gitlab:eenabulele|@eenabulele]] was approved. === 2024-03-25 === * 14:27 [[gitlab:tuukka|@tuukka]] was approved. === 2024-03-24 === * 12:24 [[gitlab:firefly|@firefly]] was approved. === 2024-03-21 === * 19:33 [[gitlab:universal-omega|@universal-omega]] was approved. === 2024-03-17 === * 10:36 [[gitlab:bisel91|@bisel91]] was approved. === 2024-03-16 === * 10:09 [[gitlab:delord|@delord]] was approved. * 00:42 [[gitlab:athulvis1|@athulvis1]] was approved. === 2024-03-15 === * 19:06 [[gitlab:ignaciorodrguez|@ignaciorodrguez]] was approved. * 08:30 [[gitlab:peachey88|@peachey88]] was approved. * 06:51 [[gitlab:derick|@derick]] was approved. === 2024-03-12 === * 15:06 [[gitlab:xiaoxiao|@xiaoxiao]] was approved. === 2024-03-06 === * 13:21 [[gitlab:desianabae1|@desianabae1]] was approved. === 2024-03-05 === * 19:21 [[gitlab:ep1c|@ep1c]] was approved. * 16:33 [[gitlab:jasmine|@jasmine]] was approved. === 2024-03-02 === * 06:42 [[gitlab:potsdamlamb|@potsdamlamb]] was approved. === 2024-02-29 === * 23:18 [[gitlab:arandomname123|@arandomname123]] was approved. * 18:03 [[gitlab:baba|@baba]] was approved. * 17:48 [[gitlab:yfdyh000|@yfdyh000]] was approved. * 03:09 [[gitlab:sds|@sds]] was approved. === 2024-02-27 === * 23:33 [[gitlab:lofhi|@lofhi]] was approved. === 2024-02-15 === * 19:45 [[gitlab:gergesshamon|@gergesshamon]] was approved. === 2024-02-14 === * 14:33 [[gitlab:philipnelson99|@philipnelson99]] was approved. === 2024-02-13 === * 13:06 [[gitlab:dringsim|@dringsim]] was approved. === 2024-02-12 === * 17:36 [[gitlab:haak|@haak]] was approved. === 2024-02-05 === * 17:33 [[gitlab:qwerfjkl|@qwerfjkl]] was approved. * 17:14 [[gitlab:ahecht|@ahecht]] was approved. === 2024-02-01 === * 09:27 [[gitlab:arinaigum|@arinaigum]] was approved. * 00:15 [[gitlab:jas42|@jas42]] was approved. * 00:15 [[gitlab:edhu|@edhu]] was approved. * 00:15 [[gitlab:marnanel|@marnanel]] was approved. * 00:15 [[gitlab:ibrahemqasim|@ibrahemqasim]] was approved. * 00:15 [[gitlab:amasotti|@amasotti]] was approved. * 00:15 [[gitlab:deni|@deni]] was approved. * 00:15 [[gitlab:cyber|@cyber]] was approved. * 00:15 [[gitlab:saroj|@saroj]] was approved. === 2024-01-29 === * 21:42 [[gitlab:rgupta|@rgupta]] was approved. === 2024-01-07 === * 09:48 [[gitlab:lutrome|@lutrome]] was approved. === 2024-01-05 === * 20:48 [[gitlab:jinoytommanjaly|@jinoytommanjaly]] was approved. * 02:51 [[gitlab:braunobruno|@braunobruno]] was approved. * 01:08 [[gitlab:amorymeltzer|@amorymeltzer]] was approved. * 01:08 [[gitlab:phi22ipus|@phi22ipus]] was approved. === 2024-01-03 === * 14:45 [[gitlab:gabina|@gabina]] was approved. === 2024-01-02 === * 13:18 [[gitlab:arthurtaylor|@arthurtaylor]] was approved. === 2023-12-23 === * 00:33 [[gitlab:aram|@aram]] was approved. === 2023-12-22 === * 16:24 [[gitlab:elpitareio|@elpitareio]] was approved. === 2023-12-21 === * 00:43 [[gitlab:bsadowski1|@bsadowski1]] was approved. * 00:43 [[gitlab:ederporto|@ederporto]] was approved. * 00:43 [[gitlab:sadraiiali|@sadraiiali]] was approved. * 00:43 [[gitlab:wasp-outis|@wasp-outis]] was approved. * 00:43 [[gitlab:bodhisattwa|@bodhisattwa]] was approved. * 00:43 [[gitlab:air7538|@air7538]] was approved. * 00:43 [[gitlab:anzx|@anzx]] was approved. * 00:43 [[gitlab:tekask1903|@tekask1903]] was approved. * 00:42 [[gitlab:kiwi-0x010c|@kiwi-0x010c]] was approved. * 00:42 [[gitlab:mpaa|@mpaa]] was approved. * 00:42 [[gitlab:kutay|@kutay]] was approved. * 00:42 [[gitlab:wattmto|@wattmto]] was approved. 9l4l7qvzj6ydgoi7149ji6hvrbwvrlq 2445275 2445262 2026-08-10T10:33:13Z Gitlabaccountapprovalbot 37332 c8263a20 was rejected. 2445275 wikitext text/x-wiki <noinclude>'''Audit log of approvals''' made by [[gitlab:gitlabaccountapprovalbot|@gitlabaccountapprovalbot]]. __NOTOC__</noinclude> === 2026-08-10 === * 10:33 "c8263a20" was rejected (pending since 2026-05-11T10:32:40.353Z). === 2026-08-09 === * 21:18 [[gitlab:iamnetx|@iamnetx]] was approved. * 20:09 "yirba" was rejected (pending since 2026-05-10T20:07:07.738Z). * 06:42 "marsam2489" was rejected (pending since 2026-05-10T06:40:46.276Z). === 2026-08-08 === * 08:15 [[gitlab:taiwaniajusto|@taiwaniajusto]] was approved. === 2026-08-07 === * 18:21 [[gitlab:gturkington|@gturkington]] was approved. * 06:36 "brianbybyby" was rejected (pending since 2026-05-08T06:36:08.459Z). === 2026-07-30 === * 20:42 "horaciocolbert" was rejected (pending since 2026-04-30T20:41:40.421Z). === 2026-07-29 === * 10:03 [[gitlab:piastu|@piastu]] was approved. * 06:33 "rafiul1" was rejected (pending since 2026-04-29T06:33:02.573Z). === 2026-07-28 === * 16:51 [[gitlab:for-each-next|@for-each-next]] was approved. * 09:21 [[gitlab:gka|@gka]] was approved. === 2026-07-27 === * 08:12 [[gitlab:cambob|@cambob]] was approved. === 2026-07-26 === * 15:39 "demansanaagmailcom" was rejected (pending since 2026-04-26T15:39:03.297Z). === 2026-07-24 === * 10:54 [[gitlab:ysogo|@ysogo]] was approved. * 09:39 [[gitlab:pankaj199|@pankaj199]] was approved. === 2026-07-23 === * 12:30 [[gitlab:fermiboson|@fermiboson]] was approved. * 09:21 [[gitlab:slashme|@slashme]] was approved. === 2026-07-22 === * 15:54 [[gitlab:lmedley|@lmedley]] was approved. * 13:24 "praveen5638" was rejected (pending since 2026-04-22T13:21:23.368Z). * 12:48 [[gitlab:cyberpower678|@cyberpower678]] was approved. * 08:30 [[gitlab:nabbegat|@nabbegat]] was approved. * 08:30 [[gitlab:plyd|@plyd]] was approved. * 07:06 "ayush8620" was rejected (pending since 2026-04-22T07:03:12.476Z). === 2026-07-21 === * 15:51 [[gitlab:panieravide|@panieravide]] was approved. * 15:33 [[gitlab:luisvilla-personal|@luisvilla-personal]] was approved. * 13:48 [[gitlab:yru|@yru]] was approved. * 13:09 [[gitlab:deevad|@deevad]] was approved. * 12:54 [[gitlab:ctdo17|@ctdo17]] was approved. * 12:36 [[gitlab:jeannenoiraud|@jeannenoiraud]] was approved. * 12:21 [[gitlab:nadiantara|@nadiantara]] was approved. * 12:18 [[gitlab:wijltcher|@wijltcher]] was approved. * 10:33 [[gitlab:nivopol|@nivopol]] was approved. * 10:30 [[gitlab:johlig|@johlig]] was approved. * 10:30 [[gitlab:majicita|@majicita]] was approved. * 10:15 [[gitlab:yongjiapeng|@yongjiapeng]] was approved. * 10:09 [[gitlab:francyskus|@francyskus]] was approved. * 09:24 [[gitlab:xanonymusx|@xanonymusx]] was approved. === 2026-07-20 === * 17:15 "leonidlednev" was rejected (pending since 2026-04-20T17:13:35.108Z). * 15:27 [[gitlab:rodrigoargenton|@rodrigoargenton]] was approved. * 05:48 "draftecho" was rejected (pending since 2026-04-20T05:48:06.953Z). === 2026-07-19 === * 12:21 [[gitlab:boivie|@boivie]] was approved. === 2026-07-18 === * 16:09 [[gitlab:pharos|@pharos]] was approved. * 15:45 [[gitlab:priyankar22|@priyankar22]] was approved. * 15:30 [[gitlab:sisyph|@sisyph]] was approved. === 2026-07-13 === * 03:45 [[gitlab:dreamyshade|@dreamyshade]] was approved. === 2026-07-12 === * 09:27 [[gitlab:smk|@smk]] was approved. === 2026-07-11 === * 14:48 "bigcereal42" was rejected (pending since 2026-04-11T14:47:40.321Z). * 12:18 "pratyushsawan" was rejected (pending since 2026-04-11T12:16:41.671Z). === 2026-07-10 === * 14:42 [[gitlab:akaza24|@akaza24]] was approved. * 12:57 [[gitlab:kormisk|@kormisk]] was approved. === 2026-07-07 === * 10:36 [[gitlab:olafjanssen|@olafjanssen]] was approved. * 06:57 "elisapoly-99" was rejected (pending since 2026-04-07T06:55:52.662Z). === 2026-07-06 === * 11:57 "ma3rouf" was rejected (pending since 2026-04-06T11:56:30.978Z). === 2026-07-02 === * 15:48 [[gitlab:tekneos|@tekneos]] was approved. === 2026-07-01 === * 15:12 [[gitlab:mugurolevy|@mugurolevy]] was approved. * 14:15 [[gitlab:vadymts1|@vadymts1]] was approved. * 09:57 "mugurolevy" was rejected (pending since 2026-04-01T09:55:19.175Z). === 2026-06-30 === * 14:27 "shivangisharma" was rejected (pending since 2026-03-31T14:26:44.932Z). === 2026-06-29 === * 19:03 [[gitlab:thisismattmiller|@thisismattmiller]] was approved. === 2026-06-28 === * 14:51 "nkwenuinadine" was rejected (pending since 2026-03-29T14:48:32.735Z). * 14:03 "vaishnavikumbhar" was rejected (pending since 2026-03-29T14:01:30.604Z). * 13:03 "stepmay" was rejected (pending since 2026-03-29T13:01:59.905Z). * 06:51 "swallroth" was rejected (pending since 2026-03-29T06:49:54.838Z). === 2026-06-26 === * 09:39 [[gitlab:lakshita28|@lakshita28]] was approved. * 07:30 [[gitlab:reeti|@reeti]] was approved. * 07:30 [[gitlab:anushka10patel|@anushka10patel]] was approved. * 07:30 "samsaesque" was rejected (pending since 2026-03-27T07:29:57.279Z). * 05:51 [[gitlab:arpithhhaaa|@arpithhhaaa]] was approved. * 05:51 [[gitlab:govindlaltl|@govindlaltl]] was approved. === 2026-06-25 === * 16:09 [[gitlab:sakuraemad|@sakuraemad]] was approved. * 07:00 "kdh8219" was rejected (pending since 2026-03-26T06:58:05.415Z). === 2026-06-24 === * 11:54 [[gitlab:sanskardubeydev|@sanskardubeydev]] was approved. * 10:09 "tanmay789q" was rejected (pending since 2026-03-25T10:07:54.602Z). === 2026-06-22 === * 19:57 [[gitlab:gouvernathor|@gouvernathor]] was approved. * 16:45 [[gitlab:lucasbelo|@lucasbelo]] was approved. * 07:15 "jason2000-cpu" was rejected (pending since 2026-03-23T07:14:09.184Z). === 2026-06-21 === * 13:18 [[gitlab:egonw|@egonw]] was approved. === 2026-06-20 === * 10:21 [[gitlab:tways2017|@tways2017]] was approved. === 2026-06-19 === * 16:06 "wilsonwang2026" was rejected (pending since 2026-03-20T16:06:05.511Z). * 04:12 [[gitlab:claudio|@claudio]] was approved. === 2026-06-18 === * 14:21 "royiswariii" was rejected (pending since 2026-03-19T14:19:16.896Z). * 13:06 [[gitlab:laurabarluzzi|@laurabarluzzi]] was approved. === 2026-06-17 === * 11:24 "adinathq8x" was rejected (pending since 2026-03-18T11:22:50.098Z). * 09:45 "nathanveritas" was rejected (pending since 2026-03-18T09:43:51.645Z). === 2026-06-15 === * 22:39 [[gitlab:mohammadhijjawi|@mohammadhijjawi]] was approved. * 14:24 "enlisar" was rejected (pending since 2026-03-16T14:23:00.109Z). * 14:06 "ayaan" was rejected (pending since 2026-03-16T14:03:31.071Z). * 10:54 "kwametech" was rejected (pending since 2026-03-16T10:54:11.083Z). === 2026-06-14 === * 17:45 [[gitlab:surajseth520|@surajseth520]] was approved. * 07:24 "malahimhaseeb" was rejected (pending since 2026-03-15T07:21:57.748Z). === 2026-06-11 === * 11:48 [[gitlab:cadddr|@cadddr]] was approved. * 11:18 "wikipiggy" was rejected (pending since 2026-03-12T11:16:09.335Z). * 07:15 [[gitlab:vesihiisi|@vesihiisi]] was approved. === 2026-06-10 === * 07:03 [[gitlab:dmiranda|@dmiranda]] was approved. === 2026-06-09 === * 14:21 [[gitlab:linkgenetic|@linkgenetic]] was approved. * 14:03 [[gitlab:sjones-ctr|@sjones-ctr]] was approved. * 12:51 [[gitlab:ekrem|@ekrem]] was approved. === 2026-06-08 === * 17:48 "jmprax" was rejected (pending since 2026-03-09T17:46:38.807Z). * 16:15 [[gitlab:ahonc|@ahonc]] was approved. * 12:57 [[gitlab:rainmonger|@rainmonger]] was approved. === 2026-06-07 === * 23:03 "shadowthewuff" was rejected (pending since 2026-03-08T23:00:53.442Z). * 11:45 "wiki-pavan" was rejected (pending since 2026-03-08T11:45:11.116Z). * 02:39 [[gitlab:launchpad|@launchpad]] was approved. === 2026-06-06 === * 14:54 "unicord" was rejected (pending since 2026-03-07T14:52:04.992Z). * 12:48 "chien" was rejected (pending since 2026-03-07T12:48:11.669Z). === 2026-06-04 === * 14:33 "only-vikas" was rejected (pending since 2026-03-05T14:32:09.186Z). === 2026-06-03 === * 15:00 [[gitlab:anafibnshahibul|@anafibnshahibul]] was approved. === 2026-06-02 === * 21:21 "mgagat" was rejected (pending since 2026-03-03T21:18:37.223Z). * 13:57 "prasunaenumarthy" was rejected (pending since 2026-03-03T13:57:14.847Z). * 05:48 [[gitlab:tmoney|@tmoney]] was approved. === 2026-06-01 === * 14:57 "vikram2101" was rejected (pending since 2026-03-02T14:54:26.550Z). * 12:03 "watshell" was rejected (pending since 2026-03-02T12:03:09.329Z). === 2026-05-29 === * 12:48 "mounikapotladurthi" was rejected (pending since 2026-02-27T12:45:38.609Z). === 2026-05-27 === * 20:00 "vinitha" was rejected (pending since 2026-02-25T19:58:43.524Z). * 16:30 "codeurluce" was rejected (pending since 2026-02-25T16:28:53.973Z). * 14:33 [[gitlab:thilio|@thilio]] was approved. === 2026-05-26 === * 12:09 "charisad" was rejected (pending since 2026-02-24T12:07:21.881Z). === 2026-05-25 === * 22:54 "ddshelto" was rejected (pending since 2026-02-23T22:52:44.427Z). * 19:51 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z). * 19:48 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z). === 2026-05-24 === * 18:45 "jiyagupta-cs" was rejected (pending since 2026-02-22T18:43:33.176Z). === 2026-05-23 === * 13:09 [[gitlab:gauthammohanraj|@gauthammohanraj]] was approved. * 04:21 [[gitlab:staraction|@staraction]] was approved. === 2026-05-22 === * 19:03 "i-horich" was rejected (pending since 2026-02-20T19:00:43.519Z). * 01:48 "50323233" was rejected (pending since 2026-02-20T01:48:05.555Z). === 2026-05-21 === * 18:51 "kartikeyg0104" was rejected (pending since 2026-02-19T18:48:39.707Z). * 16:27 [[gitlab:renovatebot|@renovatebot]] was approved. * 16:06 [[gitlab:gkm563|@gkm563]] was approved. === 2026-05-20 === * 01:21 "beedellrokejulianlockhart" was rejected (pending since 2026-02-18T01:19:13.284Z). === 2026-05-18 === * 23:18 "wladek92" was rejected (pending since 2026-02-16T23:16:22.939Z). * 16:36 [[gitlab:effeietsanders|@effeietsanders]] was approved. === 2026-05-14 === * 21:00 [[gitlab:nehemienathan|@nehemienathan]] was approved. === 2026-05-13 === * 10:51 "ssssaaaa" was rejected (pending since 2026-02-11T10:50:36.975Z). === 2026-05-12 === * 18:06 [[gitlab:psubhashish|@psubhashish]] was approved. * 08:12 "khan" was rejected (pending since 2026-02-10T08:11:48.776Z). * 04:27 "galaxysh" was rejected (pending since 2026-02-10T04:24:59.440Z). === 2026-05-11 === * 12:18 "peterxy12" was rejected (pending since 2026-02-09T12:18:01.982Z). === 2026-05-10 === * 11:09 "yalihupokn" was rejected (pending since 2026-02-08T11:06:51.336Z). * 05:12 "wobadha" was rejected (pending since 2026-02-08T05:11:00.569Z). === 2026-05-09 === * 13:45 "bwiki" was rejected (pending since 2026-02-07T13:43:38.177Z). === 2026-05-08 === * 09:24 [[gitlab:cwilliams|@cwilliams]] was approved. === 2026-05-07 === * 14:15 "rehankhan78" was rejected (pending since 2026-02-05T14:13:37.754Z). === 2026-05-06 === * 11:24 "ari" was rejected (pending since 2026-02-04T11:24:11.760Z). * 08:09 [[gitlab:neriah|@neriah]] was approved. * 06:27 [[gitlab:status401|@status401]] was approved. === 2026-05-03 === * 09:54 [[gitlab:anilk|@anilk]] was approved. === 2026-05-02 === * 17:54 [[gitlab:sweil|@sweil]] was approved. * 17:00 [[gitlab:aoppo|@aoppo]] was approved. === 2026-05-01 === * 21:18 [[gitlab:dawalda|@dawalda]] was approved. === 2026-04-30 === * 21:42 "merohibine" was rejected (pending since 2026-01-29T21:40:00.756Z). * 20:54 [[gitlab:tfmorris|@tfmorris]] was approved. * 17:33 [[gitlab:uyen|@uyen]] was approved. * 07:39 [[gitlab:mahveotm|@mahveotm]] was approved. * 06:36 [[gitlab:leo321|@leo321]] was approved. === 2026-04-29 === * 02:27 [[gitlab:dw31415|@dw31415]] was approved. === 2026-04-28 === * 23:09 [[gitlab:dtorsani|@dtorsani]] was approved. === 2026-04-27 === * 23:42 [[gitlab:quinlan|@quinlan]] was approved. * 05:00 [[gitlab:matthewyeager|@matthewyeager]] was approved. === 2026-04-26 === * 17:36 "kuba-hajnej" was rejected (pending since 2026-01-25T17:33:32.467Z). * 13:03 "jklamo" was rejected (pending since 2026-01-25T13:02:22.936Z). === 2026-04-25 === * 20:24 [[gitlab:maldaxura|@maldaxura]] was approved. * 14:33 [[gitlab:sirtobi|@sirtobi]] was approved. * 04:18 "ice5678" was rejected (pending since 2026-01-24T04:15:30.008Z). === 2026-04-24 === * 22:06 [[gitlab:arcstur|@arcstur]] was approved. === 2026-04-22 === * 23:06 "dtorsani" was rejected (pending since 2026-01-21T23:03:25.843Z). * 22:18 [[gitlab:egezort|@egezort]] was approved. * 16:45 "nexpectarpit" was rejected (pending since 2026-01-21T16:43:21.045Z). === 2026-04-20 === * 19:15 "fitch" was rejected (pending since 2026-01-19T19:12:35.644Z). === 2026-04-19 === * 02:54 [[gitlab:neoact|@neoact]] was approved. === 2026-04-18 === * 07:06 [[gitlab:kockaadmiralac|@kockaadmiralac]] was approved. === 2026-04-17 === * 13:42 "liselot" was rejected (pending since 2026-01-16T13:39:41.909Z). === 2026-04-15 === * 17:03 "lahari" was rejected (pending since 2026-01-14T17:02:06.275Z). === 2026-04-14 === * 13:00 "surajseth520" was rejected (pending since 2026-01-13T12:59:45.906Z). * 04:51 [[gitlab:canley|@canley]] was approved. * 01:03 "bshizzle" was rejected (pending since 2026-01-13T01:00:48.120Z). === 2026-04-13 === * 15:30 [[gitlab:passimacopoulos|@passimacopoulos]] was approved. === 2026-04-11 === * 12:30 "krithash" was rejected (pending since 2026-01-10T12:27:24.731Z). === 2026-04-10 === * 15:30 "raunak1709" was rejected (pending since 2026-01-09T15:29:10.901Z). === 2026-04-07 === * 17:03 [[gitlab:supnabla|@supnabla]] was approved. === 2026-04-06 === * 20:00 [[gitlab:laerdon|@laerdon]] was approved. * 19:21 [[gitlab:ljq3|@ljq3]] was approved. === 2026-04-04 === * 11:06 "mixcc" was rejected (pending since 2026-01-03T11:03:33.922Z). === 2026-04-02 === * 05:30 [[gitlab:mbh1|@mbh1]] was approved. === 2026-04-01 === * 18:21 "yuvrajpatil17" was rejected (pending since 2025-12-31T18:20:27.991Z). * 12:12 [[gitlab:amorii0|@amorii0]] was approved. === 2026-03-31 === * 11:00 "krrishsehgal" was rejected (pending since 2025-12-30T11:00:16.384Z). === 2026-03-30 === * 15:36 [[gitlab:atsuko|@atsuko]] was approved. === 2026-03-29 === * 11:36 [[gitlab:giftcup|@giftcup]] was approved. === 2026-03-28 === * 14:51 [[gitlab:janeeva1|@janeeva1]] was approved. === 2026-03-26 === * 13:36 [[gitlab:saiphani02|@saiphani02]] was approved. * 11:48 [[gitlab:valerioboz-wmch|@valerioboz-wmch]] was approved. === 2026-03-25 === * 09:45 "quansi" was rejected (pending since 2025-12-24T09:42:13.451Z). * 02:18 [[gitlab:viztor|@viztor]] was approved. === 2026-03-24 === * 23:18 [[gitlab:maryyann|@maryyann]] was approved. * 23:01 [[gitlab:codenamenoreste|@codenamenoreste]] was approved. * 13:36 [[gitlab:marc-maillard-wmse|@marc-maillard-wmse]] was approved. * 07:39 "fred2675" was rejected (pending since 2025-12-23T07:39:11.380Z). === 2026-03-23 === * 14:51 [[gitlab:komla|@komla]] was approved. * 05:51 "lunachuck43" was rejected (pending since 2025-12-22T05:50:17.862Z). * 04:06 "reza110011" was rejected (pending since 2025-12-22T04:05:25.117Z). === 2026-03-20 === * 21:54 "mertgor" was rejected (pending since 2025-12-19T21:51:51.419Z). * 20:57 "autanmahmah" was rejected (pending since 2025-12-19T20:54:51.678Z). * 09:57 [[gitlab:nethahussain|@nethahussain]] was approved. * 09:27 [[gitlab:piewriter|@piewriter]] was approved. * 08:15 [[gitlab:dondersmooi|@dondersmooi]] was approved. === 2026-03-19 === * 21:03 "sayvhior" was rejected (pending since 2025-12-18T21:02:31.699Z). === 2026-03-18 === * 20:15 [[gitlab:martinmystere|@martinmystere]] was approved. === 2026-03-17 === * 02:51 "louperivois" was rejected (pending since 2025-12-16T02:50:48.197Z). === 2026-03-16 === * 12:54 "mokayaj857" was rejected (pending since 2025-12-15T12:53:39.015Z). * 06:18 "roamer15" was rejected (pending since 2025-12-15T06:16:38.042Z). === 2026-03-14 === * 11:12 "umaramuhammad" was rejected (pending since 2025-12-13T11:10:44.004Z). * 09:33 "akuma19" was rejected (pending since 2025-12-13T09:31:39.044Z). * 07:06 [[gitlab:syunsyunminmin|@syunsyunminmin]] was approved. === 2026-03-12 === * 20:24 [[gitlab:11wb|@11wb]] was approved. * 09:54 [[gitlab:bcxfu75k|@bcxfu75k]] was approved. === 2026-03-10 === * 09:12 [[gitlab:viktoriahillerudwmse|@viktoriahillerudwmse]] was approved. === 2026-03-06 === * 08:09 "vazhayilnewone" was rejected (pending since 2025-12-05T08:07:02.184Z). === 2026-03-04 === * 20:54 [[gitlab:elphie|@elphie]] was approved. * 11:39 "ronaldahmed" was rejected (pending since 2025-12-03T11:37:47.492Z). * 02:12 "ltslw" was rejected (pending since 2025-12-03T02:11:52.040Z). === 2026-03-02 === * 19:21 "dlopez350" was rejected (pending since 2025-12-01T19:20:38.918Z). * 18:15 [[gitlab:lsandergreen|@lsandergreen]] was approved. === 2026-03-01 === * 10:51 [[gitlab:clintacc|@clintacc]] was approved. === 2026-02-28 === * 09:24 "cardboardlamp" was rejected (pending since 2025-11-29T09:22:03.947Z). * 08:18 "wiki-pavan" was rejected (pending since 2025-11-29T08:16:24.184Z). === 2026-02-27 === * 20:45 "thisisrick25" was rejected (pending since 2025-11-28T20:42:24.454Z). === 2026-02-26 === * 13:57 "chuiimuiiofc" was rejected (pending since 2025-11-27T13:57:02.794Z). * 13:54 "steffpro" was rejected (pending since 2025-11-27T13:52:10.859Z). === 2026-02-25 === * 21:24 "abubakarhabibudayyabu" was rejected (pending since 2025-11-26T21:22:37.776Z). === 2026-02-24 === * 05:00 "playboi" was rejected (pending since 2025-11-25T05:00:30.762Z). === 2026-02-23 === * 14:00 "alph65" was rejected (pending since 2025-11-24T13:59:00.797Z). * 12:33 [[gitlab:robertsky|@robertsky]] was approved. === 2026-02-22 === * 00:30 "hp8p" was rejected (pending since 2025-11-23T00:29:24.741Z). === 2026-02-19 === * 16:45 "clayjar" was rejected (pending since 2025-11-20T16:44:48.380Z). === 2026-02-18 === * 22:18 "nexus" was rejected (pending since 2025-11-19T22:16:48.818Z). * 12:00 "bernsteinnn" was rejected (pending since 2025-11-19T11:59:04.427Z). === 2026-02-17 === * 11:36 "jason2000-cpu" was rejected (pending since 2025-11-18T11:34:00.314Z). === 2026-02-16 === * 14:54 "smaurya" was rejected (pending since 2025-11-17T14:52:06.906Z). === 2026-02-15 === * 16:51 "kra-79" was rejected (pending since 2025-11-16T16:50:41.375Z). === 2026-02-14 === * 15:15 [[gitlab:mess|@mess]] was approved. === 2026-02-13 === * 13:57 "sopalsuemae957" was rejected (pending since 2025-11-14T13:55:16.921Z). * 13:30 [[gitlab:wyslijp16-toolforge|@wyslijp16-toolforge]] was approved. === 2026-02-12 === * 16:30 "kristinagligoric" was rejected (pending since 2025-11-13T16:29:21.646Z). * 03:33 [[gitlab:anyehansen|@anyehansen]] was approved. * 02:21 [[gitlab:thejoyfultentmaker|@thejoyfultentmaker]] was approved. === 2026-02-10 === * 13:18 [[gitlab:db111|@db111]] was approved. === 2026-02-09 === * 19:06 "squirrel289" was rejected (pending since 2025-11-10T19:04:27.831Z). === 2026-02-06 === * 20:54 [[gitlab:gillux|@gillux]] was approved. * 09:09 [[gitlab:lih|@lih]] was approved. === 2026-01-31 === * 16:21 [[gitlab:taxonbot1|@taxonbot1]] was approved. === 2026-01-28 === * 14:30 [[gitlab:ademola|@ademola]] was approved. * 10:51 "watshell" was rejected (pending since 2025-10-29T10:51:01.521Z). === 2026-01-26 === * 23:06 "tavaresgmg" was rejected (pending since 2025-10-27T23:04:42.140Z). === 2026-01-25 === * 06:03 "cata" was rejected (pending since 2025-10-26T06:01:26.155Z). === 2026-01-24 === * 21:15 [[gitlab:wiegels|@wiegels]] was approved. * 06:30 [[gitlab:blaquans|@blaquans]] was approved. === 2026-01-23 === * 16:27 [[gitlab:lerickson|@lerickson]] was approved. * 10:15 "fran0035g" was rejected (pending since 2025-10-24T10:12:17.732Z). === 2026-01-22 === * 21:00 "hacksyn" was rejected (pending since 2025-10-23T20:59:15.982Z). === 2026-01-21 === * 17:30 [[gitlab:otcenas11|@otcenas11]] was approved. === 2026-01-19 === * 21:48 [[gitlab:amdrel|@amdrel]] was approved. * 04:36 "rayalexa" was rejected (pending since 2025-10-20T04:35:02.094Z). === 2026-01-18 === * 15:45 "somya" was rejected (pending since 2025-10-19T15:43:43.701Z). * 06:54 "sergg001" was rejected (pending since 2025-10-19T06:54:12.296Z). === 2026-01-16 === * 11:57 "zeejohsy" was rejected (pending since 2025-10-17T11:56:22.372Z). * 04:45 "rocky25" was rejected (pending since 2025-10-17T04:43:33.180Z). === 2026-01-15 === * 16:39 "tiisu" was rejected (pending since 2025-10-16T16:37:18.438Z). * 12:00 "noahalorwu" was rejected (pending since 2025-10-16T11:58:26.133Z). * 10:39 "prjayaiuedu" was rejected (pending since 2025-10-16T10:37:16.947Z). === 2026-01-13 === * 17:21 [[gitlab:lwilson-ctr|@lwilson-ctr]] was approved. === 2026-01-12 === * 17:03 "stagietechs" was rejected (pending since 2025-10-13T17:02:25.281Z). === 2026-01-10 === * 19:06 "keerthisr" was rejected (pending since 2025-10-11T19:05:01.758Z). === 2026-01-09 === * 20:36 "lightb" was rejected (pending since 2025-10-10T20:34:20.264Z). === 2026-01-08 === * 19:42 [[gitlab:tbodt|@tbodt]] was approved. * 13:57 [[gitlab:martynranyard|@martynranyard]] was approved. === 2026-01-07 === * 17:48 [[gitlab:santanuwiki25|@santanuwiki25]] was approved. * 14:27 "dipanshu" was rejected (pending since 2025-10-08T14:26:10.794Z). * 12:30 "adeolaadesina" was rejected (pending since 2025-10-08T12:29:49.592Z). * 09:21 "tony-kamande" was rejected (pending since 2025-10-08T09:20:28.421Z). * 06:18 "hninwuttyi" was rejected (pending since 2025-10-08T06:17:28.006Z). * 05:09 "andume" was rejected (pending since 2025-10-08T05:07:18.582Z). * 02:00 "mosope" was rejected (pending since 2025-10-08T01:59:54.800Z). * 01:15 [[gitlab:tungstalite|@tungstalite]] was approved. === 2026-01-06 === * 18:24 "leerensucher" was rejected (pending since 2025-10-07T18:21:41.253Z). * 14:54 "leonidlednev" was rejected (pending since 2025-10-07T14:53:07.273Z). * 12:57 "alexandre-tingaud" was rejected (pending since 2025-10-07T12:54:27.206Z). === 2026-01-04 === * 21:33 [[gitlab:matr1x-101|@matr1x-101]] was approved. * 15:18 "makjr" was rejected (pending since 2025-10-05T15:16:31.558Z). * 14:09 "dakshq" was rejected (pending since 2025-10-05T14:08:40.608Z). === 2026-01-03 === * 20:42 [[gitlab:apehitkey|@apehitkey]] was approved. * 18:00 [[gitlab:jeremyb|@jeremyb]] was approved. * 14:09 [[gitlab:twelephant|@twelephant]] was approved. === 2026-01-01 === * 11:30 "shellstanislav" was rejected (pending since 2025-10-02T11:29:10.150Z). === 2025-12-30 === * 19:51 "camilojdiaz" was rejected (pending since 2025-09-30T19:49:24.913Z). === 2025-12-29 === * 16:03 "zied" was rejected (pending since 2025-09-29T16:01:30.415Z). * 08:18 "rahulsidpradhan" was rejected (pending since 2025-09-29T08:17:02.849Z). === 2025-12-26 === * 09:48 "thembo42" was rejected (pending since 2025-09-26T09:45:15.033Z). === 2025-12-25 === * 14:03 "196936074751" was rejected (pending since 2025-09-25T14:02:31.367Z). === 2025-12-23 === * 16:21 "ngarnsworthy" was rejected (pending since 2025-09-23T16:20:41.211Z). === 2025-12-22 === * 12:39 "aza555" was rejected (pending since 2025-09-22T12:38:02.622Z). === 2025-12-20 === * 23:45 "saph" was rejected (pending since 2025-09-20T23:45:01.222Z). === 2025-12-19 === * 10:15 "vladdymoses" was rejected (pending since 2025-09-19T10:15:00.999Z). * 07:15 "dirtylittlepoobah" was rejected (pending since 2025-09-19T07:13:55.537Z). === 2025-12-18 === * 16:24 [[gitlab:guyfawcus|@guyfawcus]] was approved. === 2025-12-17 === * 21:39 [[gitlab:holdyourhorses|@holdyourhorses]] was approved. * 18:30 "prudencia" was rejected (pending since 2025-09-17T18:27:18.860Z). * 02:24 "lottie" was rejected (pending since 2025-09-17T02:21:21.744Z). === 2025-12-16 === * 09:39 [[gitlab:melcatherine|@melcatherine]] was approved. * 08:54 [[gitlab:leila237|@leila237]] was approved. === 2025-12-15 === * 18:27 [[gitlab:royalsailor|@royalsailor]] was approved. * 09:39 [[gitlab:olaf8940|@olaf8940]] was approved. * 09:39 "brianbybyby" was rejected (pending since 2025-09-15T09:37:45.430Z). === 2025-12-14 === * 20:21 [[gitlab:essa237|@essa237]] was approved. * 16:42 [[gitlab:bovimacoco|@bovimacoco]] was approved. === 2025-12-13 === * 21:54 "mmns21" was rejected (pending since 2025-09-13T21:52:24.017Z). * 20:33 "bugcrawler" was rejected (pending since 2025-09-13T20:31:09.211Z). === 2025-12-12 === * 14:39 "ruvchoudhary" was rejected (pending since 2025-09-12T14:36:16.167Z). * 06:54 "rezadress" was rejected (pending since 2025-09-12T06:52:21.749Z). === 2025-12-10 === * 17:30 [[gitlab:itsmoon|@itsmoon]] was approved. === 2025-12-09 === * 15:42 [[gitlab:mercy-o|@mercy-o]] was approved. === 2025-12-06 === * 16:45 "jacquesradjabu" was rejected (pending since 2025-09-06T16:45:17.969Z). * 11:27 [[gitlab:ikhitron|@ikhitron]] was approved. === 2025-12-01 === * 08:12 "halconmilenario21" was rejected (pending since 2025-09-01T08:12:10.262Z). === 2025-11-30 === * 21:06 [[gitlab:habs|@habs]] was approved. === 2025-11-29 === * 16:36 "bovimacoco" was rejected (pending since 2025-08-30T16:34:39.712Z). * 00:45 [[gitlab:jjpmaster|@jjpmaster]] was approved. === 2025-11-24 === * 10:30 "alph65" was rejected (pending since 2025-08-25T10:28:40.957Z). * 02:24 [[gitlab:yaron|@yaron]] was approved. === 2025-11-20 === * 16:06 "clayjar" was rejected (pending since 2025-08-21T16:04:54.450Z). === 2025-11-17 === * 21:09 [[gitlab:ankita97531|@ankita97531]] was approved. === 2025-11-16 === * 14:15 "commanderkefir" was rejected (pending since 2025-08-17T14:13:14.791Z). * 08:21 "rehankhan78" was rejected (pending since 2025-08-17T08:19:44.896Z). === 2025-11-15 === * 14:36 "cyberscribe" was rejected (pending since 2025-08-16T14:34:27.230Z). === 2025-11-13 === * 04:21 "waddie96" was rejected (pending since 2025-08-14T04:19:27.461Z). === 2025-11-11 === * 06:42 [[gitlab:seanhoyland|@seanhoyland]] was approved. === 2025-11-10 === * 00:06 [[gitlab:jaredblumer|@jaredblumer]] was approved. === 2025-11-09 === * 22:36 "heinxiety" was rejected (pending since 2025-08-10T22:33:12.041Z). === 2025-11-07 === * 22:00 [[gitlab:forzagreen|@forzagreen]] was approved. === 2025-11-06 === * 16:57 [[gitlab:rsilvola|@rsilvola]] was approved. === 2025-11-04 === * 21:24 [[gitlab:devdoingdev|@devdoingdev]] was approved. === 2025-11-03 === * 17:48 "joewaleed98" was rejected (pending since 2025-08-04T17:46:12.191Z). === 2025-11-01 === * 18:00 "eliasempresas" was rejected (pending since 2025-08-02T17:58:04.412Z). === 2025-10-31 === * 18:51 [[gitlab:chaoticenby|@chaoticenby]] was approved. * 04:33 "3ch310n" was rejected (pending since 2025-08-01T04:32:21.982Z). === 2025-10-30 === * 10:03 [[gitlab:tausheefhassan|@tausheefhassan]] was approved. === 2025-10-29 === * 14:54 "theap" was rejected (pending since 2025-07-30T14:52:12.066Z). === 2025-10-28 === * 06:06 [[gitlab:tanbiruzzaman|@tanbiruzzaman]] was approved. === 2025-10-27 === * 07:51 [[gitlab:jmoore111|@jmoore111]] was approved. === 2025-10-25 === * 21:09 [[gitlab:valor|@valor]] was approved. * 21:03 [[gitlab:booksmurf|@booksmurf]] was approved. * 02:48 "mystyc1" was rejected (pending since 2025-07-26T02:46:19.373Z). === 2025-10-24 === * 05:12 "aadarshmahesh" was rejected (pending since 2025-07-25T05:09:38.264Z). === 2025-10-22 === * 20:54 [[gitlab:janewanga|@janewanga]] was approved. * 17:27 "abeljeevan" was rejected (pending since 2025-07-23T17:26:46.884Z). * 16:12 "shrimpnaur" was rejected (pending since 2025-07-23T16:10:37.864Z). === 2025-10-21 === * 18:51 "jrmuizel" was rejected (pending since 2025-07-22T18:50:07.315Z). * 09:33 [[gitlab:dpogorzelski|@dpogorzelski]] was approved. === 2025-10-17 === * 13:21 [[gitlab:blegodwin|@blegodwin]] was approved. === 2025-10-16 === * 14:51 [[gitlab:bahago|@bahago]] was approved. * 14:12 "harikrishna0005" was rejected (pending since 2025-07-17T14:10:48.385Z). * 14:09 "gauthammohanraj" was rejected (pending since 2025-07-17T14:08:47.643Z). === 2025-10-15 === * 13:48 [[gitlab:adwivedii|@adwivedii]] was approved. * 13:18 [[gitlab:kimbrenekakande|@kimbrenekakande]] was approved. * 13:03 "childmnajennifer" was rejected (pending since 2025-07-16T13:01:50.236Z). * 05:06 "vssb4214" was rejected (pending since 2025-07-16T05:05:33.985Z). === 2025-10-14 === * 19:39 [[gitlab:afanyulionel|@afanyulionel]] was approved. * 15:33 [[gitlab:sadrettin|@sadrettin]] was approved. * 14:18 [[gitlab:tmwyk|@tmwyk]] was approved. * 08:42 "yasu0796" was rejected (pending since 2025-07-15T08:41:26.453Z). === 2025-10-13 === * 16:09 [[gitlab:atlas0007|@atlas0007]] was approved. === 2025-10-11 === * 17:42 [[gitlab:techwizzie|@techwizzie]] was approved. === 2025-10-10 === * 19:03 [[gitlab:miiswom|@miiswom]] was approved. * 16:06 [[gitlab:ninatakang|@ninatakang]] was approved. === 2025-10-09 === * 15:42 [[gitlab:jaykaneki|@jaykaneki]] was approved. * 14:21 [[gitlab:lebogang|@lebogang]] was approved. * 14:15 [[gitlab:kimondorose|@kimondorose]] was approved. * 13:48 [[gitlab:joyakinyi|@joyakinyi]] was approved. * 13:48 [[gitlab:dikshyashahi|@dikshyashahi]] was approved. * 13:45 [[gitlab:obediobadiah|@obediobadiah]] was approved. * 13:45 [[gitlab:system625|@system625]] was approved. * 13:45 [[gitlab:rolalove|@rolalove]] was approved. * 13:39 [[gitlab:olatundeawo|@olatundeawo]] was approved. * 13:36 [[gitlab:danielchristlight|@danielchristlight]] was approved. * 13:36 [[gitlab:dipanshu1223|@dipanshu1223]] was approved. * 13:36 [[gitlab:aradhya|@aradhya]] was approved. * 09:57 "bognd" was rejected (pending since 2025-07-10T09:55:48.661Z). === 2025-10-08 === * 23:36 [[gitlab:sopzy|@sopzy]] was approved. * 23:03 [[gitlab:oluwatumininu|@oluwatumininu]] was approved. * 19:39 [[gitlab:levon003|@levon003]] was approved. * 15:24 [[gitlab:ritika-bhambri11|@ritika-bhambri11]] was approved. * 13:45 [[gitlab:anbanguyen|@anbanguyen]] was approved. * 13:36 [[gitlab:chumzine|@chumzine]] was approved. * 13:27 [[gitlab:shr0x-ya|@shr0x-ya]] was approved. * 12:45 [[gitlab:nurahwakili|@nurahwakili]] was approved. * 03:42 "nazhiba" was rejected (pending since 2025-07-09T03:40:12.625Z). * 02:12 "mafennel" was rejected (pending since 2025-07-09T02:11:40.598Z). === 2025-10-07 === * 22:54 [[gitlab:olusegunfaj|@olusegunfaj]] was approved. * 21:30 [[gitlab:rona|@rona]] was approved. * 21:09 [[gitlab:sandijigs|@sandijigs]] was approved. * 13:36 "xisbajao" was rejected (pending since 2025-07-08T13:33:35.018Z). * 01:36 "areczek94" was rejected (pending since 2025-07-08T01:35:40.633Z). === 2025-10-06 === * 19:21 "wmcarter2017" was rejected (pending since 2025-07-07T19:21:12.899Z). === 2025-10-05 === * 14:15 "meetmendapara" was rejected (pending since 2025-07-06T14:14:16.726Z). === 2025-10-04 === * 20:51 "nftbaee" was rejected (pending since 2025-07-05T20:50:57.688Z). === 2025-10-03 === * 06:12 [[gitlab:javiermonton|@javiermonton]] was approved. === 2025-10-02 === * 20:15 "talaqalotaibipmp" was rejected (pending since 2025-07-03T20:13:05.164Z). === 2025-10-01 === * 10:54 "bjensen" was rejected (pending since 2025-07-02T10:53:46.574Z). * 02:45 "kowal1984" was rejected (pending since 2025-07-02T02:44:56.946Z). === 2025-09-30 === * 21:21 [[gitlab:kavaljeetsingh|@kavaljeetsingh]] was approved. * 00:24 "adium" was rejected (pending since 2025-07-01T00:23:43.807Z). === 2025-09-28 === * 08:54 [[gitlab:pexerik|@pexerik]] was approved. === 2025-09-27 === * 13:57 [[gitlab:rubahhitamvukova|@rubahhitamvukova]] was approved. === 2025-09-26 === * 16:57 "algorithmic" was rejected (pending since 2025-06-27T16:56:17.480Z). * 13:54 [[gitlab:shadabgdg|@shadabgdg]] was approved. * 13:12 [[gitlab:spushpit|@spushpit]] was approved. === 2025-09-20 === * 14:06 "bwiki" was rejected (pending since 2025-06-21T13:59:14.749Z). === 2025-09-16 === * 05:39 [[gitlab:deepchirp|@deepchirp]] was approved. === 2025-09-15 === * 22:00 [[gitlab:noisk8|@noisk8]] was approved. * 11:03 "ahonc" was rejected (pending since 2025-06-16T11:00:54.843Z). === 2025-09-13 === * 18:24 "a-ssh22" was rejected (pending since 2025-06-14T18:23:33.937Z). * 12:36 [[gitlab:rajashreetalukdar|@rajashreetalukdar]] was approved. * 00:45 [[gitlab:sumitsurai|@sumitsurai]] was approved. === 2025-09-12 === * 17:12 [[gitlab:suyash23|@suyash23]] was approved. * 00:46 "remotetravel" was rejected (pending since 2025-06-13T00:44:08.171Z). === 2025-09-10 === * 21:09 "jancborchardt" was rejected (pending since 2025-06-11T21:06:30.759Z). === 2025-09-09 === * 17:03 [[gitlab:vwf|@vwf]] was approved. * 06:36 [[gitlab:cactusisme|@cactusisme]] was approved. === 2025-09-08 === * 18:09 "birushandegeya" was rejected (pending since 2025-06-09T18:08:00.087Z). * 16:27 "ngarnsworthy" was rejected (pending since 2025-06-09T16:24:37.213Z). * 12:33 "zolgoyo" was rejected (pending since 2025-06-09T12:31:34.199Z). === 2025-09-06 === * 23:09 [[gitlab:jaishsingh913|@jaishsingh913]] was approved. === 2025-09-05 === * 21:45 [[gitlab:sakshi2|@sakshi2]] was approved. * 20:42 "abdukhaliq1" was rejected (pending since 2025-06-06T20:40:42.023Z). * 14:27 "beubsamy" was rejected (pending since 2025-06-06T14:27:06.781Z). === 2025-09-04 === * 23:27 "sdhehua" was rejected (pending since 2025-06-05T23:24:45.777Z). * 19:00 [[gitlab:perry|@perry]] was approved. * 11:24 "saintwolf" was rejected (pending since 2025-06-05T11:21:20.176Z). === 2025-09-02 === * 05:48 [[gitlab:aliu|@aliu]] was approved. === 2025-08-29 === * 13:30 "kksurendran066" was rejected (pending since 2025-05-30T13:27:48.755Z). === 2025-08-28 === * 22:18 "tauraamuix" was rejected (pending since 2025-05-29T22:16:08.228Z). === 2025-08-26 === * 19:03 [[gitlab:dikkulah|@dikkulah]] was approved. === 2025-08-22 === * 23:51 [[gitlab:khoroshun_mike|@khoroshun_mike]] was approved. === 2025-08-21 === * 07:39 [[gitlab:yuka|@yuka]] was approved. === 2025-08-19 === * 07:48 [[gitlab:zhaofjx|@zhaofjx]] was approved. === 2025-08-17 === * 14:27 "madhan13k" was rejected (pending since 2025-05-18T14:26:08.973Z). === 2025-08-15 === * 10:15 "mohammed_abukhadra" was rejected (pending since 2025-05-16T10:14:48.403Z). === 2025-08-11 === * 11:48 "hmmyesbro" was rejected (pending since 2025-05-12T11:45:24.350Z). === 2025-08-10 === * 13:15 [[gitlab:dactyl|@dactyl]] was approved. === 2025-08-09 === * 04:39 "xxxx100000" was rejected (pending since 2025-05-10T04:37:44.949Z). === 2025-08-08 === * 14:33 [[gitlab:josefanthony|@josefanthony]] was approved. === 2025-08-07 === * 23:42 [[gitlab:robins7|@robins7]] was approved. * 21:42 [[gitlab:pols12|@pols12]] was approved. * 17:15 "sbronson" was rejected (pending since 2025-05-08T17:15:08.834Z). * 14:57 [[gitlab:alvindulle|@alvindulle]] was approved. * 14:45 [[gitlab:xentos|@xentos]] was approved. * 06:27 "jamesboste" was rejected (pending since 2025-05-08T06:25:14.793Z). * 03:57 "ysun" was rejected (pending since 2025-05-08T03:55:07.348Z). === 2025-08-06 === * 21:51 "pols12" was rejected (pending since 2025-05-07T21:49:13.598Z). * 01:51 "okeamah" was rejected (pending since 2025-05-07T01:48:50.114Z). === 2025-08-05 === * 09:15 "mobashir-2013" was rejected (pending since 2025-05-06T09:14:24.069Z). === 2025-08-01 === * 08:00 "douginamug" was rejected (pending since 2025-05-02T07:57:38.317Z). === 2025-07-31 === * 02:30 [[gitlab:ads|@ads]] was approved. === 2025-07-27 === * 13:15 "mrico2703" was rejected (pending since 2025-04-27T13:13:12.346Z). * 10:17 [[gitlab:josephfrancis12|@josephfrancis12]] was approved. * 10:17 [[gitlab:fuzzew|@fuzzew]] was approved. * 05:57 [[gitlab:biscuitbobby|@biscuitbobby]] was approved. * 05:48 [[gitlab:ecoholic|@ecoholic]] was approved. === 2025-07-26 === * 11:48 [[gitlab:chimnayyyy|@chimnayyyy]] was approved. * 11:48 [[gitlab:alwinalbert|@alwinalbert]] was approved. * 11:48 [[gitlab:hridyakk|@hridyakk]] was approved. * 11:45 [[gitlab:gaurigupta21|@gaurigupta21]] was approved. * 11:45 [[gitlab:binetaa|@binetaa]] was approved. * 10:21 [[gitlab:jyothikat22|@jyothikat22]] was approved. * 10:21 [[gitlab:zobotrombie|@zobotrombie]] was approved. * 10:21 [[gitlab:flykrth|@flykrth]] was approved. * 10:21 [[gitlab:mehrinshamim|@mehrinshamim]] was approved. * 10:21 [[gitlab:aadhi13|@aadhi13]] was approved. * 10:21 [[gitlab:malavikam05|@malavikam05]] was approved. * 10:18 [[gitlab:nf609|@nf609]] was approved. * 05:48 [[gitlab:nazalnihad|@nazalnihad]] was approved. * 05:48 [[gitlab:naveen28204280|@naveen28204280]] was approved. === 2025-07-25 === * 09:49 [[gitlab:kasyap9|@kasyap9]] was approved. * 09:30 [[gitlab:swayamagrahari|@swayamagrahari]] was approved. === 2025-07-24 === * 19:36 [[gitlab:madutgn|@madutgn]] was approved. === 2025-07-23 === * 20:09 [[gitlab:somerandomdeveloper|@somerandomdeveloper]] was approved. === 2025-07-22 === * 00:15 [[gitlab:iagoqnsi|@iagoqnsi]] was approved. === 2025-07-21 === * 17:30 [[gitlab:asadiqui|@asadiqui]] was approved. * 16:39 [[gitlab:tryvix1509|@tryvix1509]] was approved. * 04:27 [[gitlab:damian|@damian]] was approved. === 2025-07-20 === * 09:42 "mike-khoroshun" was rejected (pending since 2025-04-20T09:42:22.732Z). === 2025-07-17 === * 17:57 [[gitlab:haroldkrabs|@haroldkrabs]] was approved. * 13:45 [[gitlab:envlh|@envlh]] was approved. === 2025-07-14 === * 10:24 [[gitlab:missguru|@missguru]] was approved. * 00:57 "clarfonthey" was rejected (pending since 2025-04-14T00:56:32.626Z). === 2025-07-13 === * 01:01 [[gitlab:l235|@l235]] was approved. === 2025-07-11 === * 03:06 "rodavlas" was rejected (pending since 2025-04-11T03:05:45.590Z). === 2025-07-06 === * 00:09 "lakasa" was rejected (pending since 2025-04-06T00:06:28.469Z). === 2025-07-05 === * 21:54 "ctrlzvi" was rejected (pending since 2025-04-05T21:54:12.542Z). * 14:30 "aminualiyu" was rejected (pending since 2025-04-05T14:27:22.617Z). === 2025-07-04 === * 03:15 [[gitlab:galstar|@galstar]] was approved. === 2025-07-02 === * 11:27 "vicolas11" was rejected (pending since 2025-04-02T11:25:12.682Z). === 2025-06-29 === * 23:12 "naomi723" was rejected (pending since 2025-03-30T23:09:24.630Z). === 2025-06-28 === * 16:21 "mudeh2372" was rejected (pending since 2025-03-29T16:18:27.057Z). === 2025-06-27 === * 23:18 "rony143" was rejected (pending since 2025-03-28T23:16:13.671Z). * 22:21 [[gitlab:rluts|@rluts]] was approved. === 2025-06-26 === * 13:54 "creativegurus" was rejected (pending since 2025-03-27T13:52:41.706Z). === 2025-06-24 === * 17:42 [[gitlab:devjadiya|@devjadiya]] was approved. * 14:00 "dominic-r" was rejected (pending since 2025-03-25T14:00:07.307Z). === 2025-06-21 === * 00:48 [[gitlab:vriaa|@vriaa]] was approved. === 2025-06-18 === * 15:21 "ayushkhati1" was rejected (pending since 2025-03-19T15:18:50.062Z). === 2025-06-17 === * 20:45 "chiomavero" was rejected (pending since 2025-03-18T20:44:13.967Z). * 00:27 [[gitlab:eggroll97|@eggroll97]] was approved. === 2025-06-14 === * 20:57 "volvox" was rejected (pending since 2025-03-15T20:56:34.018Z). === 2025-06-13 === * 16:09 [[gitlab:supergrey|@supergrey]] was approved. * 11:03 "chqaz" was rejected (pending since 2025-03-14T11:01:09.600Z). * 10:24 [[gitlab:slong-wmf|@slong-wmf]] was approved. * 10:15 "hearvox" was rejected (pending since 2025-03-14T10:13:13.112Z). === 2025-06-12 === * 15:18 "jlam" was rejected (pending since 2025-03-13T15:17:54.099Z). === 2025-06-09 === * 20:48 "dipanjansengupta" was rejected (pending since 2025-03-10T20:48:03.545Z). * 19:27 [[gitlab:reggycelly|@reggycelly]] was approved. * 14:51 "arendpieter" was rejected (pending since 2025-03-10T14:51:01.445Z). * 13:21 [[gitlab:greenreaper|@greenreaper]] was approved. * 09:33 [[gitlab:mmta|@mmta]] was approved. * 08:03 "a-ssh22" was rejected (pending since 2025-03-10T08:03:08.111Z). === 2025-06-08 === * 21:06 "mm-episodenlistedlvaupdater" was rejected (pending since 2025-03-09T21:04:06.323Z). === 2025-06-06 === * 11:06 [[gitlab:olea|@olea]] was approved. === 2025-06-05 === * 20:33 [[gitlab:encodedwp|@encodedwp]] was approved. * 15:00 [[gitlab:toluayo|@toluayo]] was approved. * 13:51 [[gitlab:arnold_lup|@arnold_lup]] was approved. * 11:54 "sdhehua" was rejected (pending since 2025-03-06T11:51:48.241Z). === 2025-06-03 === * 21:27 [[gitlab:wewakey|@wewakey]] was approved. * 12:36 "hunsimon2" was rejected (pending since 2025-03-04T12:34:56.520Z). * 11:54 "hunsimon" was rejected (pending since 2025-03-04T11:53:54.652Z). === 2025-06-02 === * 12:01 [[gitlab:jaimedes|@jaimedes]] was approved. === 2025-05-30 === * 18:00 "sathvik9105" was rejected (pending since 2025-02-28T17:59:42.867Z). * 11:21 [[gitlab:tonythomas01|@tonythomas01]] was approved. * 10:06 [[gitlab:gpsleo|@gpsleo]] was approved. === 2025-05-29 === * 22:12 [[gitlab:codynguyen1116|@codynguyen1116]] was approved. === 2025-05-28 === * 02:57 [[gitlab:saper|@saper]] was approved. === 2025-05-27 === * 21:06 [[gitlab:mohammed_qays|@mohammed_qays]] was approved. * 15:33 "satanluimm" was rejected (pending since 2025-02-25T15:32:48.101Z). === 2025-05-26 === * 23:57 "seyedali220" was rejected (pending since 2025-02-24T23:56:17.621Z). === 2025-05-21 === * 11:12 [[gitlab:guilherme|@guilherme]] was approved. === 2025-05-19 === * 13:24 [[gitlab:emojiwiki|@emojiwiki]] was approved. === 2025-05-18 === * 00:00 "xidme" was rejected (pending since 2025-02-15T23:58:56.796Z). === 2025-05-17 === * 02:39 "kdh8219" was rejected (pending since 2025-02-15T02:36:32.237Z). === 2025-05-16 === * 15:09 [[gitlab:maxbinderwmf|@maxbinderwmf]] was approved. === 2025-05-15 === * 04:30 "inspectorzer0" was rejected (pending since 2025-02-13T04:27:33.179Z). === 2025-05-14 === * 17:42 [[gitlab:llugo|@llugo]] was approved. === 2025-05-13 === * 20:18 "mmta" was rejected (pending since 2025-02-11T20:17:23.407Z). === 2025-05-11 === * 20:51 "jad" was rejected (pending since 2025-02-09T20:49:07.333Z). * 17:54 "nishchalsundan" was rejected (pending since 2025-02-09T17:52:25.761Z). * 16:39 "mohammed_abukhadra" was rejected (pending since 2025-02-09T16:39:03.730Z). === 2025-05-09 === * 09:12 [[gitlab:sirchanmp|@sirchanmp]] was approved. === 2025-05-08 === * 08:18 [[gitlab:mengeditch|@mengeditch]] was approved. === 2025-05-07 === * 03:45 "xluffy" was rejected (pending since 2025-02-05T03:45:14.181Z). === 2025-05-06 === * 16:54 "punhaniabhishek" was rejected (pending since 2025-02-04T16:53:50.758Z). * 09:36 [[gitlab:bmartinezcalvo|@bmartinezcalvo]] was approved. === 2025-05-02 === * 12:24 [[gitlab:tohaomg|@tohaomg]] was approved. * 11:48 [[gitlab:mavrikant|@mavrikant]] was approved. * 11:45 [[gitlab:daanvr|@daanvr]] was approved. === 2025-05-01 === * 09:09 "mjoerg" was rejected (pending since 2025-01-30T09:09:04.204Z). === 2025-04-30 === * 23:06 "sanskardubey" was rejected (pending since 2025-01-29T23:03:25.489Z). === 2025-04-29 === * 16:00 "geyslein" was rejected (pending since 2025-01-28T16:00:01.510Z). === 2025-04-26 === * 09:30 "anjali9027" was rejected (pending since 2025-01-25T09:28:07.064Z). === 2025-04-25 === * 18:00 "salahhazaa" was rejected (pending since 2025-01-24T17:58:30.030Z). * 15:15 [[gitlab:yiming|@yiming]] was approved. * 02:06 "mrchanmp" was rejected (pending since 2025-01-24T02:03:58.308Z). === 2025-04-23 === * 17:03 "rj2904" was rejected (pending since 2025-01-22T17:03:11.207Z). * 14:21 "nischay33" was rejected (pending since 2025-01-22T14:19:21.081Z). === 2025-04-22 === * 19:27 "dj80" was rejected (pending since 2025-01-21T19:25:28.498Z). * 14:30 [[gitlab:kaimamin|@kaimamin]] was approved. * 09:57 "debo" was rejected (pending since 2025-01-21T09:54:47.955Z). === 2025-04-21 === * 12:24 "unshell" was rejected (pending since 2025-01-20T12:21:59.686Z). === 2025-04-18 === * 15:06 [[gitlab:spartanarbinger|@spartanarbinger]] was approved. === 2025-04-16 === * 03:09 "dewey" was rejected (pending since 2025-01-15T03:06:17.488Z). === 2025-04-15 === * 19:45 "emdadul" was rejected (pending since 2025-01-14T19:42:29.285Z). === 2025-04-14 === * 06:45 [[gitlab:bcampbell804|@bcampbell804]] was approved. === 2025-04-11 === * 06:27 [[gitlab:jvanderhoop|@jvanderhoop]] was approved. === 2025-04-10 === * 04:12 "bhai420" was rejected (pending since 2025-01-09T04:10:29.430Z). === 2025-04-09 === * 05:03 "austinvarshney" was rejected (pending since 2025-01-08T05:02:34.175Z). === 2025-04-06 === * 15:36 [[gitlab:elph|@elph]] was approved. === 2025-04-02 === * 10:33 [[gitlab:ozge|@ozge]] was approved. === 2025-03-31 === * 20:15 "demandkey" was rejected (pending since 2024-12-30T20:14:23.096Z). * 15:18 [[gitlab:danyya|@danyya]] was approved. === 2025-03-28 === * 15:54 [[gitlab:rutsavi09|@rutsavi09]] was approved. * 15:54 [[gitlab:ilanen1|@ilanen1]] was approved. === 2025-03-25 === * 19:27 [[gitlab:irfo|@irfo]] was approved. * 11:54 [[gitlab:kmontalva-wmf|@kmontalva-wmf]] was approved. * 04:33 [[gitlab:paul26|@paul26]] was approved. * 04:18 "as1100k" was rejected (pending since 2024-12-24T04:18:06.813Z). === 2025-03-24 === * 11:33 "amzadkhankk" was rejected (pending since 2024-12-23T11:33:14.176Z). === 2025-03-23 === * 12:24 "wolfdo" was rejected (pending since 2024-12-22T12:23:35.056Z). === 2025-03-22 === * 09:45 [[gitlab:fjmustak|@fjmustak]] was approved. === 2025-03-20 === * 18:42 "sathishkokila" was rejected (pending since 2024-12-19T18:39:35.161Z). * 17:03 [[gitlab:alien4444|@alien4444]] was approved. * 15:27 [[gitlab:davidcoronel|@davidcoronel]] was approved. === 2025-03-19 === * 22:57 [[gitlab:r1f4t|@r1f4t]] was approved. * 19:03 "daniel24ps" was rejected (pending since 2024-12-18T19:00:21.249Z). * 14:18 [[gitlab:beepbooppenguin|@beepbooppenguin]] was approved. === 2025-03-18 === * 17:48 "rahulkundu1209" was rejected (pending since 2024-12-17T17:46:41.936Z). * 08:15 "kirtisikka972" was rejected (pending since 2024-12-17T08:13:25.487Z). === 2025-03-15 === * 13:30 "tulspal_sidhu" was rejected (pending since 2024-12-14T13:29:10.606Z). * 01:39 "peacedeadc" was rejected (pending since 2024-12-14T01:37:36.579Z). === 2025-03-14 === * 03:51 [[gitlab:chuckthebuck|@chuckthebuck]] was approved. * 02:33 "yxngtrtxll" was rejected (pending since 2024-12-13T02:31:51.658Z). === 2025-03-13 === * 14:36 [[gitlab:iccander|@iccander]] was approved. === 2025-03-12 === * 23:21 "jokerchic36" was rejected (pending since 2024-12-11T23:21:00.670Z). * 15:30 [[gitlab:naomi|@naomi]] was approved. * 15:27 [[gitlab:cobi|@cobi]] was approved. === 2025-03-11 === * 12:42 "mohitvermaxx" was rejected (pending since 2024-12-10T12:40:56.967Z). === 2025-03-10 === * 16:51 [[gitlab:nanona15dobato|@nanona15dobato]] was approved. === 2025-03-09 === * 22:39 [[gitlab:jonkolbert|@jonkolbert]] was approved. * 20:45 [[gitlab:urbanecmtest2|@urbanecmtest2]] was approved. === 2025-03-07 === * 16:54 [[gitlab:hswan|@hswan]] was approved. * 14:42 [[gitlab:atitkov|@atitkov]] was approved. * 00:42 [[gitlab:infrastruktur|@infrastruktur]] was approved. === 2025-03-06 === * 17:21 "johnmann" was rejected (pending since 2024-12-05T17:19:24.995Z). === 2025-03-05 === * 07:33 [[gitlab:monx9494|@monx9494]] was approved. === 2025-03-02 === * 21:21 "paul26" was rejected (pending since 2024-12-01T21:20:19.681Z). === 2025-03-01 === * 19:15 [[gitlab:izno|@izno]] was approved. * 12:45 [[gitlab:nyerho|@nyerho]] was approved. === 2025-02-28 === * 18:27 [[gitlab:chuckonwumelu|@chuckonwumelu]] was approved. * 13:09 "ashwinpraveengo" was rejected (pending since 2024-11-29T13:07:47.240Z). * 00:18 "eduardoaugusto" was rejected (pending since 2024-11-29T00:17:43.372Z). === 2025-02-27 === * 20:39 "volkanurl" was rejected (pending since 2024-11-28T20:37:18.101Z). === 2025-02-24 === * 21:15 [[gitlab:feeglgeef|@feeglgeef]] was approved. * 20:18 [[gitlab:piaanalysis2|@piaanalysis2]] was approved. * 19:06 [[gitlab:dhardy|@dhardy]] was approved. === 2025-02-22 === * 19:27 [[gitlab:owuh|@owuh]] was approved. === 2025-02-19 === * 16:06 [[gitlab:artemkloko|@artemkloko]] was approved. * 13:03 [[gitlab:jgafnea|@jgafnea]] was approved. === 2025-02-17 === * 16:33 [[gitlab:asmartkitten|@asmartkitten]] was approved. === 2025-02-16 === * 19:12 "gaurigupta21" was rejected (pending since 2024-11-17T19:11:07.416Z). === 2025-02-15 === * 01:18 [[gitlab:mediawiki-quickstart-ci|@mediawiki-quickstart-ci]] was approved. === 2025-02-14 === * 15:21 "nathanbnm" was rejected (pending since 2024-11-15T15:18:19.632Z). === 2025-02-13 === * 16:45 [[gitlab:priyanshuchahal|@priyanshuchahal]] was approved. * 16:42 [[gitlab:ajhalili2006|@ajhalili2006]] was approved. === 2025-02-12 === * 23:21 "monkeypatch999" was rejected (pending since 2024-11-13T23:20:38.398Z). * 06:36 [[gitlab:jainlakshita28|@jainlakshita28]] was approved. === 2025-02-11 === * 19:27 [[gitlab:matthewsm2|@matthewsm2]] was approved. === 2025-02-09 === * 16:15 "mohammed_abukhadra" was rejected (pending since 2024-11-10T16:15:18.361Z). === 2025-02-07 === * 21:33 "brennan" was rejected (pending since 2024-11-08T21:31:07.351Z). === 2025-02-06 === * 08:24 "mmta" was rejected (pending since 2024-11-07T08:22:36.724Z). * 06:21 [[gitlab:bunnypranav|@bunnypranav]] was approved. === 2025-02-05 === * 22:39 "chrissteinchen" was rejected (pending since 2024-11-06T22:38:16.673Z). === 2025-02-03 === * 07:45 "edriiic" was rejected (pending since 2024-11-04T07:44:46.849Z). * 01:12 "geppy" was rejected (pending since 2024-11-04T01:10:48.710Z). === 2025-02-02 === * 13:18 "funa-enpitu" was rejected (pending since 2024-11-03T13:15:46.065Z). === 2025-01-31 === * 23:42 "nfontes" was rejected (pending since 2024-11-01T23:39:41.755Z). * 22:51 "sbronson" was rejected (pending since 2024-11-01T22:50:31.871Z). * 00:42 [[gitlab:farid|@farid]] was approved. === 2025-01-27 === * 08:15 [[gitlab:eliza189|@eliza189]] was approved. === 2025-01-25 === * 09:51 [[gitlab:pamputt|@pamputt]] was approved. === 2025-01-23 === * 14:30 [[gitlab:lubianat|@lubianat]] was approved. * 11:45 [[gitlab:bootsa|@bootsa]] was approved. === 2025-01-21 === * 05:09 "niko" was rejected (pending since 2024-07-21T16:10:01.377Z). * 05:09 "thawizkid369777" was rejected (pending since 2024-07-18T17:42:44.493Z). * 05:09 "sarthaksingh2" was rejected (pending since 2024-07-10T11:31:30.470Z). * 05:09 "shriyakt" was rejected (pending since 2024-07-06T04:54:10.248Z). * 05:09 "akshaya" was rejected (pending since 2024-07-06T04:04:51.488Z). * 05:09 "alaka03aj" was rejected (pending since 2024-07-05T18:01:54.876Z). * 05:09 "sulochanaviji-5049" was rejected (pending since 2024-07-01T05:58:00.427Z). * 05:09 "nayanjnath" was rejected (pending since 2024-07-01T02:51:57.405Z). * 05:09 "sd44" was rejected (pending since 2024-06-30T04:28:51.436Z). * 05:09 "metavalent" was rejected (pending since 2024-06-29T01:37:14.210Z). * 05:09 "wicloudx" was rejected (pending since 2024-06-28T11:51:23.335Z). * 05:09 "debo" was rejected (pending since 2024-06-28T01:44:59.845Z). * 05:09 "bwiki" was rejected (pending since 2024-06-23T14:15:38.032Z). * 05:09 "toprak" was rejected (pending since 2024-06-23T11:35:50.819Z). * 05:09 "iristeller" was rejected (pending since 2024-06-14T20:53:48.959Z). * 05:09 "jcolvin" was rejected (pending since 2024-06-12T17:29:01.238Z). * 05:09 "kalyan" was rejected (pending since 2024-06-07T07:52:46.993Z). * 05:09 "bluecrystal" was rejected (pending since 2024-06-06T19:16:20.107Z). * 05:09 "iftttrohit" was rejected (pending since 2024-06-04T12:08:50.818Z). * 05:09 "pogpotato" was rejected (pending since 2024-06-03T17:58:21.684Z). * 05:09 "cptlausebaer" was rejected (pending since 2024-05-31T18:53:27.692Z). * 05:09 "hdevine825" was rejected (pending since 2024-05-31T17:04:18.279Z). * 05:09 "anaghaa18" was rejected (pending since 2024-05-25T19:14:31.803Z). * 05:09 "atharvanair04" was rejected (pending since 2024-05-25T14:24:52.825Z). * 05:09 "anasvemmully" was rejected (pending since 2024-05-25T06:10:27.261Z). * 05:09 "abhinavmohandas" was rejected (pending since 2024-05-25T06:05:24.825Z). * 05:09 "kksurendran06" was rejected (pending since 2024-05-25T06:04:38.082Z). * 05:09 "albertmarshall8896" was rejected (pending since 2024-05-23T09:32:05.462Z). * 05:09 "akellison" was rejected (pending since 2024-05-17T02:07:24.229Z). * 05:09 "mainowill" was rejected (pending since 2024-04-16T23:30:33.881Z). * 05:09 "bzhqc" was rejected (pending since 2024-04-16T19:50:38.676Z). * 05:09 "safan41" was rejected (pending since 2024-04-16T03:34:48.942Z). * 05:09 "mgagat" was rejected (pending since 2024-04-16T03:21:51.764Z). * 05:09 "okeamah" was rejected (pending since 2024-04-16T02:49:00.143Z). * 05:09 "xuhao61" was rejected (pending since 2024-04-15T23:45:09.083Z). * 04:47 "cybel" was rejected (pending since 2024-04-15T06:46:35.791Z). === 2025-01-20 === * 14:33 [[gitlab:your1|@your1]] was approved. === 2025-01-18 === * 10:09 [[gitlab:galrach600|@galrach600]] was approved. * 02:51 [[gitlab:blankeclair|@blankeclair]] was approved. === 2025-01-17 === * 13:57 [[gitlab:dsantamaria|@dsantamaria]] was approved. === 2025-01-15 === * 17:12 [[gitlab:smartse|@smartse]] was approved. === 2025-01-14 === * 17:03 [[gitlab:naorleizer|@naorleizer]] was approved. === 2025-01-13 === * 02:45 [[gitlab:wolf20482|@wolf20482]] was approved. === 2025-01-12 === * 17:45 [[gitlab:tamzin|@tamzin]] was approved. === 2025-01-11 === * 15:24 [[gitlab:bargioni|@bargioni]] was approved. * 14:30 [[gitlab:salelya|@salelya]] was approved. * 10:15 [[gitlab:malakatshy|@malakatshy]] was approved. * 05:21 [[gitlab:newmcpee|@newmcpee]] was approved. === 2025-01-09 === * 15:30 [[gitlab:gkyziridis|@gkyziridis]] was approved. === 2025-01-08 === * 16:21 [[gitlab:ukrface|@ukrface]] was approved. === 2024-12-28 === * 03:27 [[gitlab:twonum|@twonum]] was approved. === 2024-12-25 === * 06:09 [[gitlab:harsv567|@harsv567]] was approved. === 2024-12-21 === * 11:24 [[gitlab:amutha2002|@amutha2002]] was approved. === 2024-12-20 === * 19:51 [[gitlab:hridyeshgupta|@hridyeshgupta]] was approved. * 10:00 [[gitlab:ro-shines|@ro-shines]] was approved. * 08:09 [[gitlab:kesharwaniarpita|@kesharwaniarpita]] was approved. === 2024-12-18 === * 14:45 [[gitlab:soylacarli|@soylacarli]] was approved. === 2024-12-16 === * 20:33 [[gitlab:aleyasiddika1|@aleyasiddika1]] was approved. === 2024-12-15 === * 07:33 [[gitlab:abhishek02bhardwaj|@abhishek02bhardwaj]] was approved. === 2024-12-13 === * 13:18 [[gitlab:ashmitabathre204|@ashmitabathre204]] was approved. === 2024-12-10 === * 06:39 [[gitlab:ginaan|@ginaan]] was approved. === 2024-12-09 === * 05:45 [[gitlab:kallinavya|@kallinavya]] was approved. * 00:54 [[gitlab:viserion-7|@viserion-7]] was approved. === 2024-12-08 === * 17:27 [[gitlab:wargo|@wargo]] was approved. === 2024-12-05 === * 11:15 [[gitlab:ranjithraj|@ranjithraj]] was approved. === 2024-12-02 === * 21:21 [[gitlab:a930913|@a930913]] was approved. === 2024-12-01 === * 02:39 [[gitlab:kingchristlike1|@kingchristlike1]] was approved. === 2024-11-21 === * 13:45 [[gitlab:sascha|@sascha]] was approved. === 2024-11-19 === * 16:36 [[gitlab:jly|@jly]] was approved. === 2024-11-15 === * 02:54 [[gitlab:danielyepezgarces|@danielyepezgarces]] was approved. === 2024-11-14 === * 14:15 [[gitlab:stimoroll|@stimoroll]] was approved. === 2024-11-09 === * 17:15 [[gitlab:f4udeveloper|@f4udeveloper]] was approved. === 2024-11-07 === * 19:15 [[gitlab:zulf|@zulf]] was approved. * 05:33 [[gitlab:hassanamin|@hassanamin]] was approved. === 2024-11-06 === * 19:39 [[gitlab:daniuu|@daniuu]] was approved. * 00:18 [[gitlab:rlopez-wmf|@rlopez-wmf]] was approved. === 2024-10-09 === * 14:45 [[gitlab:jtweed|@jtweed]] was approved. * 10:24 [[gitlab:ifrahkh|@ifrahkh]] was approved. * 09:06 [[gitlab:wikibayer|@wikibayer]] was approved. === 2024-10-06 === * 10:27 [[gitlab:keerthan16|@keerthan16]] was approved. === 2024-10-04 === * 07:45 [[gitlab:hakimi97|@hakimi97]] was approved. === 2024-09-30 === * 07:39 [[gitlab:ninjastrikers|@ninjastrikers]] was approved. === 2024-09-28 === * 17:30 [[gitlab:webrunner95|@webrunner95]] was approved. === 2024-09-18 === * 21:39 [[gitlab:elliottetzkorn|@elliottetzkorn]] was approved. === 2024-09-14 === * 22:06 [[gitlab:humptydumpty|@humptydumpty]] was approved. === 2024-09-06 === * 08:48 [[gitlab:mickabarber|@mickabarber]] was approved. === 2024-08-27 === * 17:36 [[gitlab:edgars|@edgars]] was approved. === 2024-08-22 === * 09:18 [[gitlab:antonkokhwmde|@antonkokhwmde]] was approved. === 2024-08-14 === * 19:21 [[gitlab:jfk|@jfk]] was approved. === 2024-08-13 === * 17:57 [[gitlab:daxserver|@daxserver]] was approved. === 2024-08-11 === * 09:57 [[gitlab:pauliesnug|@pauliesnug]] was approved. === 2024-08-10 === * 08:42 [[gitlab:ashig|@ashig]] was approved. === 2024-08-09 === * 14:09 [[gitlab:masssly|@masssly]] was approved. === 2024-08-05 === * 22:15 [[gitlab:mrtortue|@mrtortue]] was approved. === 2024-08-02 === * 16:21 [[gitlab:dsantini|@dsantini]] was approved. === 2024-07-31 === * 11:54 [[gitlab:cptviraj|@cptviraj]] was approved. === 2024-07-30 === * 19:09 [[gitlab:iniquity|@iniquity]] was approved. * 10:00 [[gitlab:collins|@collins]] was approved. === 2024-07-27 === * 15:57 [[gitlab:songnguxyz|@songnguxyz]] was approved. === 2024-07-25 === * 12:36 [[gitlab:mszabo|@mszabo]] was approved. * 09:21 [[gitlab:agarwalmahima|@agarwalmahima]] was approved. === 2024-07-24 === * 08:05 [[gitlab:dragoniez|@dragoniez]] was approved. === 2024-07-23 === * 06:54 [[gitlab:mirji|@mirji]] was approved. === 2024-07-16 === * 10:00 [[gitlab:lakejason0|@lakejason0]] was approved. === 2024-07-12 === * 11:33 [[gitlab:cn|@cn]] was approved. * 08:12 [[gitlab:unchampignon|@unchampignon]] was approved. === 2024-07-07 === * 17:12 [[gitlab:agamyasamuel|@agamyasamuel]] was approved. * 05:24 [[gitlab:kuldeepburjbhalaike|@kuldeepburjbhalaike]] was approved. === 2024-07-06 === * 11:18 [[gitlab:dibya|@dibya]] was approved. * 04:54 [[gitlab:sarthakparashar|@sarthakparashar]] was approved. === 2024-07-05 === * 18:15 [[gitlab:vanshikarathi|@vanshikarathi]] was approved. === 2024-07-02 === * 19:00 [[gitlab:ebrahim|@ebrahim]] was approved. === 2024-07-01 === * 20:12 [[gitlab:rockingpenny4|@rockingpenny4]] was approved. * 18:15 [[gitlab:balajijagadesh|@balajijagadesh]] was approved. === 2024-06-30 === * 18:24 [[gitlab:hrideshmg|@hrideshmg]] was approved. * 07:18 [[gitlab:chanakyakumardas|@chanakyakumardas]] was approved. * 06:30 [[gitlab:rihaan180|@rihaan180]] was approved. === 2024-06-27 === * 17:36 [[gitlab:driedmueller|@driedmueller]] was approved. === 2024-06-19 === * 12:57 [[gitlab:audreypenven|@audreypenven]] was approved. === 2024-06-16 === * 01:18 [[gitlab:roysmith|@roysmith]] was approved. === 2024-06-08 === * 02:45 [[gitlab:jleedev|@jleedev]] was approved. === 2024-06-03 === * 13:57 [[gitlab:afeder|@afeder]] was approved. === 2024-06-01 === * 10:54 [[gitlab:florianschmitt|@florianschmitt]] was approved. === 2024-05-30 === * 16:42 [[gitlab:krlsca|@krlsca]] was approved. === 2024-05-28 === * 11:24 [[gitlab:rickijay|@rickijay]] was approved. === 2024-05-26 === * 11:18 [[gitlab:ranjithsiji|@ranjithsiji]] was approved. === 2024-05-25 === * 07:24 [[gitlab:jony|@jony]] was approved. === 2024-05-23 === * 08:45 [[gitlab:lepticed7|@lepticed7]] was approved. === 2024-05-22 === * 20:42 [[gitlab:echecs|@echecs]] was approved. === 2024-05-21 === * 13:33 [[gitlab:mbs|@mbs]] was approved. === 2024-05-19 === * 18:06 [[gitlab:ionenlaser|@ionenlaser]] was approved. === 2024-05-18 === * 23:36 [[gitlab:mdaniels5757|@mdaniels5757]] was approved. === 2024-05-17 === * 08:54 [[gitlab:grapedog|@grapedog]] was approved. === 2024-05-08 === * 19:42 [[gitlab:kelhurd|@kelhurd]] was approved. * 19:06 [[gitlab:khurd|@khurd]] was approved. === 2024-05-06 === * 19:48 [[gitlab:j3j5|@j3j5]] was approved. * 12:06 [[gitlab:tk-999|@tk-999]] was approved. === 2024-05-05 === * 22:09 [[gitlab:pppery|@pppery]] was approved. * 20:33 [[gitlab:sakretsu|@sakretsu]] was approved. * 12:12 [[gitlab:waterquark|@waterquark]] was approved. === 2024-05-04 === * 09:03 [[gitlab:multichill|@multichill]] was approved. * 07:42 [[gitlab:abaris|@abaris]] was approved. === 2024-05-03 === * 14:57 [[gitlab:maurusian|@maurusian]] was approved. === 2024-04-24 === * 05:48 [[gitlab:wolfinux|@wolfinux]] was approved. === 2024-04-23 === * 15:48 [[gitlab:dreamrimmer|@dreamrimmer]] was approved. === 2024-04-21 === * 06:51 [[gitlab:alon|@alon]] was approved. === 2024-04-17 === * 23:33 [[gitlab:derenrich|@derenrich]] was approved. === 2024-04-16 === * 17:18 [[gitlab:valcio|@valcio]] was approved. === 2024-04-14 === * 16:51 [[gitlab:wikilucas00|@wikilucas00]] was approved. === 2024-04-06 === * 12:48 [[gitlab:theprotonade|@theprotonade]] was approved. === 2024-04-02 === * 07:30 [[gitlab:bohuizhang|@bohuizhang]] was approved. === 2024-03-30 === * 13:36 [[gitlab:lpintscher|@lpintscher]] was approved. === 2024-03-26 === * 17:09 [[gitlab:eenabulele|@eenabulele]] was approved. === 2024-03-25 === * 14:27 [[gitlab:tuukka|@tuukka]] was approved. === 2024-03-24 === * 12:24 [[gitlab:firefly|@firefly]] was approved. === 2024-03-21 === * 19:33 [[gitlab:universal-omega|@universal-omega]] was approved. === 2024-03-17 === * 10:36 [[gitlab:bisel91|@bisel91]] was approved. === 2024-03-16 === * 10:09 [[gitlab:delord|@delord]] was approved. * 00:42 [[gitlab:athulvis1|@athulvis1]] was approved. === 2024-03-15 === * 19:06 [[gitlab:ignaciorodrguez|@ignaciorodrguez]] was approved. * 08:30 [[gitlab:peachey88|@peachey88]] was approved. * 06:51 [[gitlab:derick|@derick]] was approved. === 2024-03-12 === * 15:06 [[gitlab:xiaoxiao|@xiaoxiao]] was approved. === 2024-03-06 === * 13:21 [[gitlab:desianabae1|@desianabae1]] was approved. === 2024-03-05 === * 19:21 [[gitlab:ep1c|@ep1c]] was approved. * 16:33 [[gitlab:jasmine|@jasmine]] was approved. === 2024-03-02 === * 06:42 [[gitlab:potsdamlamb|@potsdamlamb]] was approved. === 2024-02-29 === * 23:18 [[gitlab:arandomname123|@arandomname123]] was approved. * 18:03 [[gitlab:baba|@baba]] was approved. * 17:48 [[gitlab:yfdyh000|@yfdyh000]] was approved. * 03:09 [[gitlab:sds|@sds]] was approved. === 2024-02-27 === * 23:33 [[gitlab:lofhi|@lofhi]] was approved. === 2024-02-15 === * 19:45 [[gitlab:gergesshamon|@gergesshamon]] was approved. === 2024-02-14 === * 14:33 [[gitlab:philipnelson99|@philipnelson99]] was approved. === 2024-02-13 === * 13:06 [[gitlab:dringsim|@dringsim]] was approved. === 2024-02-12 === * 17:36 [[gitlab:haak|@haak]] was approved. === 2024-02-05 === * 17:33 [[gitlab:qwerfjkl|@qwerfjkl]] was approved. * 17:14 [[gitlab:ahecht|@ahecht]] was approved. === 2024-02-01 === * 09:27 [[gitlab:arinaigum|@arinaigum]] was approved. * 00:15 [[gitlab:jas42|@jas42]] was approved. * 00:15 [[gitlab:edhu|@edhu]] was approved. * 00:15 [[gitlab:marnanel|@marnanel]] was approved. * 00:15 [[gitlab:ibrahemqasim|@ibrahemqasim]] was approved. * 00:15 [[gitlab:amasotti|@amasotti]] was approved. * 00:15 [[gitlab:deni|@deni]] was approved. * 00:15 [[gitlab:cyber|@cyber]] was approved. * 00:15 [[gitlab:saroj|@saroj]] was approved. === 2024-01-29 === * 21:42 [[gitlab:rgupta|@rgupta]] was approved. === 2024-01-07 === * 09:48 [[gitlab:lutrome|@lutrome]] was approved. === 2024-01-05 === * 20:48 [[gitlab:jinoytommanjaly|@jinoytommanjaly]] was approved. * 02:51 [[gitlab:braunobruno|@braunobruno]] was approved. * 01:08 [[gitlab:amorymeltzer|@amorymeltzer]] was approved. * 01:08 [[gitlab:phi22ipus|@phi22ipus]] was approved. === 2024-01-03 === * 14:45 [[gitlab:gabina|@gabina]] was approved. === 2024-01-02 === * 13:18 [[gitlab:arthurtaylor|@arthurtaylor]] was approved. === 2023-12-23 === * 00:33 [[gitlab:aram|@aram]] was approved. === 2023-12-22 === * 16:24 [[gitlab:elpitareio|@elpitareio]] was approved. === 2023-12-21 === * 00:43 [[gitlab:bsadowski1|@bsadowski1]] was approved. * 00:43 [[gitlab:ederporto|@ederporto]] was approved. * 00:43 [[gitlab:sadraiiali|@sadraiiali]] was approved. * 00:43 [[gitlab:wasp-outis|@wasp-outis]] was approved. * 00:43 [[gitlab:bodhisattwa|@bodhisattwa]] was approved. * 00:43 [[gitlab:air7538|@air7538]] was approved. * 00:43 [[gitlab:anzx|@anzx]] was approved. * 00:43 [[gitlab:tekask1903|@tekask1903]] was approved. * 00:42 [[gitlab:kiwi-0x010c|@kiwi-0x010c]] was approved. * 00:42 [[gitlab:mpaa|@mpaa]] was approved. * 00:42 [[gitlab:kutay|@kutay]] was approved. * 00:42 [[gitlab:wattmto|@wattmto]] was approved. 3cfosrfy871osluho5w1ctchrnqqmu8 Fundraising/SAL 0 459699 2445268 2444992 2026-08-10T04:21:35Z Stashbot 7414 eileen: config revision changed from e99fd679 to 4b91cb74 2445268 wikitext text/x-wiki == 2026-08-10 == * 04:21 eileen: config revision changed from {{Gerrit|e99fd679}} to {{Gerrit|4b91cb74}} == 2026-08-06 == * 14:44 larssandergreen: civicrm upgraded from {{Gerrit|56c8de1f}} to {{Gerrit|78e2ecac}} * 01:36 larssandergreen: civicrm upgraded from {{Gerrit|6408e93a}} to {{Gerrit|56c8de1f}} == 2026-08-05 == * 17:53 larssandergreen: civicrm upgraded from {{Gerrit|e229c253}} to {{Gerrit|6408e93a}} == 2026-08-04 == * 23:15 eileen: civicrm upgraded from {{Gerrit|8ed3282e}} to {{Gerrit|e229c253}} * 02:35 ejegg: fundraising civicrm upgraded from {{Gerrit|9c8ee02f}} to {{Gerrit|8ed3282e}} * 00:51 eileen: checking attributes == 2026-08-03 == * 22:27 eileen: civicrm upgraded from {{Gerrit|26bd5ab6}} to {{Gerrit|9c8ee02f}} * 22:06 eileen: SmashPig upgraded from {{Gerrit|0c2593b5}} to {{Gerrit|c89520d4}} == 2026-07-30 == * 13:53 damilare: donorwiki upgraded from {{Gerrit|d59684ed}} to {{Gerrit|0117d21d}} == 2026-07-28 == * 20:08 ejegg: fundraising civicrm upgraded from {{Gerrit|98ef4abd}} to {{Gerrit|26bd5ab6}} * 16:34 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|038018d4}} to {{Gerrit|0c2593b5}} == 2026-07-23 == * 23:50 larssandergreen: civicrm upgraded from {{Gerrit|bf93abac}} to {{Gerrit|98ef4abd}} * 00:04 eileen: civicrm upgraded from {{Gerrit|068b2e3c}} to {{Gerrit|bf93abac}} == 2026-07-22 == * 22:35 eileen: civicrm upgraded from {{Gerrit|5a04e8a9}} to {{Gerrit|068b2e3c}} * 21:38 eileen: SmashPig upgraded from {{Gerrit|6921d862}} to {{Gerrit|038018d4}} * 21:36 eileen: SmashPig upgraded from {{Gerrit|6921d862}} to {{Gerrit|038018d4}} * 01:36 eileen: civicrm upgraded from {{Gerrit|2eba4374}} to {{Gerrit|5a04e8a9}} == 2026-07-17 == * 03:44 eileen: civicrm upgraded from {{Gerrit|dd65e1b6}} to {{Gerrit|2eba4374}} * 01:25 eileen: civicrm upgraded from {{Gerrit|1c94b489}} to {{Gerrit|dd65e1b6}} == 2026-07-16 == * 11:06 laurabar: process-control config revision changed from {{Gerrit|730715e1}} to {{Gerrit|c951e730}} * 10:41 laurabar: process-control config revision changed from {{Gerrit|7dddea38}} to {{Gerrit|730715e1}} * 02:31 eileen: civicrm upgraded from {{Gerrit|2dd23034}} to {{Gerrit|1c94b489}} * 01:51 eileen: SmashPig upgraded from {{Gerrit|2e07ce09}} to {{Gerrit|6921d862}} == 2026-07-15 == * 23:48 eileen: config revision changed from {{Gerrit|8078b0b0}} to {{Gerrit|7dddea38}} * 23:12 eileen: SmashPig upgraded from {{Gerrit|37a80ec4}} to {{Gerrit|2e07ce09}} * 21:48 eileen: civicrm upgraded from {{Gerrit|d38ae4d6}} to {{Gerrit|2dd23034}} * 21:19 eileen: SmashPig upgraded from {{Gerrit|34082cca}} to {{Gerrit|37a80ec4}} * 04:58 eileen: civicrm upgraded from {{Gerrit|24d6e9a0}} to {{Gerrit|d38ae4d6}} * 04:30 eileen: SmashPig upgraded from {{Gerrit|87dfbd6b}} to {{Gerrit|34082cca}} * 03:44 cstone: payments-wiki upgraded from {{Gerrit|776f7d37}} to {{Gerrit|5482106f}} == 2026-07-14 == * 02:45 larssandergreen: tools upgraded from {{Gerrit|0c0cbb37}} to {{Gerrit|93a9d78d}} * 01:42 larssandergreen: civicrm upgraded from {{Gerrit|015e84b8}} to {{Gerrit|24d6e9a0}} == 2026-07-13 == * 22:46 larssandergreen: civicrm upgraded from {{Gerrit|171652fa}} to {{Gerrit|015e84b8}} * 19:48 eileen: civicrm upgraded from {{Gerrit|263ff140}} to {{Gerrit|171652fa}} == 2026-07-10 == * 21:17 larssandergreen: tools upgraded from {{Gerrit|afbd0f67}} to {{Gerrit|0c0cbb37}} * 20:00 cstone: civicrm upgraded from {{Gerrit|20b3dabe}} to {{Gerrit|263ff140}} == 2026-07-09 == * 05:02 eileen: civicrm upgraded from {{Gerrit|3a1a0891}} to {{Gerrit|20b3dabe}} * 03:12 eileen: civicrm upgraded from {{Gerrit|3a2952ff}} to {{Gerrit|3a1a0891}} * 02:28 eileen: civicrm upgraded from {{Gerrit|f4ff79c8}} to {{Gerrit|3a2952ff}} * 02:12 eileen: SmashPig upgraded from {{Gerrit|2042f7f7}} to {{Gerrit|87dfbd6b}} * 00:01 eileen: onfig {{Gerrit|1c59ae42}} -> {{Gerrit|fa79267c}} == 2026-07-08 == * 23:52 eileen: SmashPig upgraded from {{Gerrit|d8393666}} to {{Gerrit|2042f7f7}} * 22:15 larssandergreen: civicrm upgraded from {{Gerrit|ead61459}} to {{Gerrit|f4ff79c8}} * 22:13 larssandergreen: donorwiki upgraded from {{Gerrit|c2cdbf24}} to {{Gerrit|d59684ed}} * 19:50 jgleeson: payments-wiki upgraded from {{Gerrit|196c7bc1}} to {{Gerrit|5d2bb8e8}} * 13:46 jgleeson: payments-wiki upgraded from {{Gerrit|f34901d7}} to {{Gerrit|196c7bc1}} * 03:48 larssandergreen: civicrm upgraded from {{Gerrit|0057f9a1}} to {{Gerrit|ead61459}} * 01:12 larssandergreen: civicrm upgraded from {{Gerrit|941505c3}} to {{Gerrit|0057f9a1}} == 2026-07-07 == * 22:29 jgleeson: payments-wiki upgraded from {{Gerrit|ece97dca}} to {{Gerrit|f34901d7}} * 04:47 eileen: civicrm upgraded from {{Gerrit|0013ea5e}} to {{Gerrit|941505c3}} * 03:27 eileen: process-control * 01:35 larssandergreen: civicrm upgraded from {{Gerrit|274308a4}} to {{Gerrit|0013ea5e}} == 2026-07-02 == * 23:46 eileen: civicrm upgraded from {{Gerrit|d44caa2a}} to {{Gerrit|274308a4}} * 03:36 eileen: civicrm upgraded from {{Gerrit|7288da0a}} to {{Gerrit|d44caa2a}} * 01:54 eileen: * civicrm upgraded from {{Gerrit|13584dc4}} to {{Gerrit|7288da0a}} * 01:46 eileen: config revision changed from {{Gerrit|92ac127c}} to {{Gerrit|d0a8b49f}} == 2026-06-30 == * 21:25 eileen: SmashPig upgraded from {{Gerrit|38ffd696}} to {{Gerrit|d8393666}} * 07:01 eileen: civicrm upgraded from {{Gerrit|63606ee8}} to {{Gerrit|c2fa32a6}} == 2026-06-29 == * 23:16 danielfm: payments-wiki upgraded from {{Gerrit|7e1a422e}} to {{Gerrit|28983fa4}} * 18:54 ejegg: payments-wiki upgraded from {{Gerrit|32059801}} to {{Gerrit|7e1a422e}} * 18:32 wfan: payments-wiki upgraded from {{Gerrit|39df7f6b}} to {{Gerrit|32059801}} * 18:08 ejegg: payments-wiki upgraded from {{Gerrit|c2cdbf24}} to {{Gerrit|39df7f6b}} == 2026-06-26 == * 08:17 eileen: civicrm upgraded from {{Gerrit|a90a5451}} to {{Gerrit|ebab84c8}} == 2026-06-25 == * 13:35 jgleeson: payments-wiki upgraded from {{Gerrit|ab4fd16b}} to {{Gerrit|c2cdbf24}} == 2026-06-24 == * 17:30 ejegg: fundraising civicrm upgraded from {{Gerrit|7130d7ff}} to {{Gerrit|a90a5451}} == 2026-06-23 == * 09:29 jgleeson: payments-wiki upgraded from {{Gerrit|71cba440}} to {{Gerrit|ab4fd16b}} == 2026-06-22 == * 22:53 eileen: SmashPig upgraded from {{Gerrit|9cd51fd1}} to {{Gerrit|38ffd696}} * 22:32 eileen: config revision changed from {{Gerrit|c687d8f0}} to {{Gerrit|92ac127c}} * 22:15 eileen: civicrm upgraded from {{Gerrit|602742db}} to {{Gerrit|7130d7ff}} * 20:27 ejegg: donorwiki upgraded from {{Gerrit|1f5f40f9}} to {{Gerrit|71cba440}} * 16:35 ejegg: payments-wiki upgraded from {{Gerrit|873882a5}} to {{Gerrit|d5f4b8aa}} == 2026-06-18 == * 00:42 eileen: civicrm upgraded from {{Gerrit|646d893e}} to {{Gerrit|602742db}} == 2026-06-17 == * 01:22 ejegg: payments-wiki upgraded from {{Gerrit|1f5f40f9}} to {{Gerrit|873882a5}} == 2026-06-16 == * 20:31 ejegg: standalone (ipn listener) SmashPig upgraded from {{Gerrit|611947be}} to {{Gerrit|9cd51fd1}} * 03:02 eileen: SmashPig upgraded from {{Gerrit|8ad963c7}} to {{Gerrit|611947be}} * 01:16 eileen: config revision changed from {{Gerrit|7ca9f992}} to {{Gerrit|8334c030}} * 01:12 eileen: updated SmashPig revision {{Gerrit|e82c2c5f}} -> {{Gerrit|8ad963c7}} == 2026-06-15 == * 11:43 damilare: donorwiki upgraded from {{Gerrit|3bc70a73}} to {{Gerrit|1f5f40f9}} == 2026-06-12 == * 18:08 jgleeson: civicrm upgraded from {{Gerrit|69a60dcb}} to {{Gerrit|646d893e}} * 03:42 eileen: civicrm upgraded from {{Gerrit|0b8db587}} to {{Gerrit|69a60dcb}} == 2026-06-11 == * 18:56 jgleeson: payments-wiki upgraded from {{Gerrit|aef3d25d}} to {{Gerrit|1f5f40f9}} * 06:05 eileen: civicrm upgraded from {{Gerrit|77961b3e}} to {{Gerrit|0b8db587}} * 02:59 larssandergreen: civicrm upgraded from {{Gerrit|8f98770f}} to {{Gerrit|77961b3e}} * 02:42 eileen: civicrm upgraded from {{Gerrit|2962b28d}} to {{Gerrit|8f98770f}} * 01:30 eileen: civicrm upgraded from {{Gerrit|f46a6066}} to {{Gerrit|2962b28d}} * 00:22 eileen: civicrm upgraded from {{Gerrit|819c4ede}} to {{Gerrit|f46a6066}} == 2026-06-10 == * 16:04 wfan: civicrm upgraded from {{Gerrit|d7113e04}} to {{Gerrit|819c4ede}} * 11:36 jgleeson: SmashPig upgraded from {{Gerrit|f6b24f7f}} to {{Gerrit|e82c2c5f}} == 2026-06-09 == * 21:39 eileen: civicrm upgraded from {{Gerrit|6b61d3a5}} to {{Gerrit|d7113e04}} == 2026-06-04 == * 11:32 jgleeson: payments-wiki upgraded from {{Gerrit|3bc70a73}} to {{Gerrit|aef3d25d}} == 2026-06-03 == * 12:17 jgleeson: SmashPig upgraded from {{Gerrit|166abfbd}} to {{Gerrit|99233b18}} * 05:46 eileen: civicrm upgraded from {{Gerrit|55adc0bb}} to {{Gerrit|6b61d3a5}} * 04:28 eileen: civicrm upgraded from {{Gerrit|219cf085}} to {{Gerrit|55adc0bb}} * 03:32 eileen: civicrm upgraded from {{Gerrit|f6839b65}} to {{Gerrit|219cf085}} * 02:07 eileen: civicrm upgraded from {{Gerrit|663c0c30}} to {{Gerrit|f6839b65}} == 2026-06-02 == * 23:14 eileen: config revision changed from {{Gerrit|67231927}} to {{Gerrit|c011a0d6}} * 22:55 eileen: config revision changed from {{Gerrit|46874905}} to {{Gerrit|67231927}} == 2026-06-01 == * 19:53 eileen: civicrm upgraded from {{Gerrit|a4a885f4}} to {{Gerrit|663c0c30}} * 13:20 damilare: donorwiki upgraded from {{Gerrit|9f3734e0}} to {{Gerrit|3bc70a73}} == 2026-05-28 == * 22:09 eileen: SmashPig upgraded from {{Gerrit|252b6cfa}} to {{Gerrit|166abfbd}} * 19:44 dwisehaupt: removing apache2::mod::python from civicrm and frdev roles. * 18:43 ejegg: payments-wiki upgraded from {{Gerrit|9f3734e0}} to {{Gerrit|21185532}} * 15:13 ejegg: donorwiki upgraded from {{Gerrit|1a056dc0}} to {{Gerrit|9f3734e0}} * 15:12 ejegg: payments-wiki upgraded from {{Gerrit|1e2eb148}} to {{Gerrit|9f3734e0}} * 14:38 ejegg: civicrm upgraded from {{Gerrit|0f0567b3}} to {{Gerrit|a4a885f4}} == 2026-05-26 == * 01:05 eileen: civicrm upgraded from {{Gerrit|6b4e53a5}} to {{Gerrit|713e8508}} == 2026-05-21 == * 04:13 eileen: civicrm upgraded from {{Gerrit|cae2ab0e}} to {{Gerrit|dbafc0b4}} * 00:09 cstone: civicrm upgraded from {{Gerrit|fcbbf763}} to {{Gerrit|cae2ab0e}} == 2026-05-20 == * 22:22 eileen: SmashPig upgraded from {{Gerrit|df441f78}} to {{Gerrit|252b6cfa}} == 2026-05-19 == * 23:08 eileen: civicrm upgraded from {{Gerrit|6cff86c1}} to {{Gerrit|fcbbf763}} * 01:56 eileen: civicrm upgraded from {{Gerrit|bfa13f7f}} to {{Gerrit|6cff86c1}} == 2026-05-18 == * 20:45 eileen: civicrm upgraded from {{Gerrit|150d484e}} to {{Gerrit|bfa13f7f}} * 14:26 ejegg: payments-wiki upgraded from {{Gerrit|40a24102}} to {{Gerrit|e31ac2a9}} * 11:29 jgleeson: payments-wiki upgraded from {{Gerrit|1a056dc0}} to {{Gerrit|40a24102}} == 2026-05-15 == * 19:19 dwisehaupt: redis swap complete from frqueue1003 to frqueue1005 * 15:59 ejegg: donorwiki upgraded from {{Gerrit|26f5451a}} to {{Gerrit|1a056dc0}} * 15:35 ejegg: payments-wiki upgraded from {{Gerrit|cf9ec80b}} to {{Gerrit|1a056dc0}} * 00:02 eileen: civicrm upgraded from {{Gerrit|6d8ce7a3}} to {{Gerrit|6a2258ff}} == 2026-05-14 == * 22:22 ejegg: fundraising scheduled jobs re-enabled * 22:12 eileen: cv upgraded from {{Gerrit|f19e0961}} to {{Gerrit|b8a8dd6a}} * 22:11 ejegg: fundraising civicrm upgraded from {{Gerrit|e25fa223}} to {{Gerrit|6d8ce7a3}} * 22:09 ejegg: fundraising scheduled jobs disabled for Civi update * 19:29 ejegg: re-enabled fundraising scheduled jobs * 19:06 ejegg: fundraising civicrm upgraded from {{Gerrit|950908ec}} to {{Gerrit|e25fa223}} * 19:04 ejegg: disabled fundraising scheduled jobs for CiviCRM deployment * 16:48 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|bb833986}} to {{Gerrit|a2a3a015}} == 2026-05-12 == * 16:07 ejegg: fundraising civicrm upgraded from {{Gerrit|24ac90e7}} to {{Gerrit|60dcc28f}} == 2026-05-06 == * 12:29 eileen: config revision changed from {{Gerrit|41cfd677}} to {{Gerrit|00752f91}} * 10:56 eileen: SmashPig upgraded from {{Gerrit|4201ef56}} to {{Gerrit|bb833986}} * 09:59 eileen: civicrm upgraded from {{Gerrit|4d9c8600}} to {{Gerrit|24ac90e7}} * 09:15 eileen: civicrm upgraded from {{Gerrit|38dcf7a8}} to {{Gerrit|4d9c8600}} == 2026-05-04 == * 17:39 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|1be60746}} to {{Gerrit|4201ef56}} == 2026-05-03 == * hackathon: civicrm upgraded from {{Gerrit|0afbc8ea}} to {{Gerrit|38dcf7a8}} == 2026-05-02 == * 09:44 eileen: civicrm upgraded from {{Gerrit|7556c5c7}} to {{Gerrit|0afbc8ea}} * 09:43 eileen: SmashPig upgraded from {{Gerrit|88a1bcba}} to {{Gerrit|1be60746}} == 2026-05-01 == * 16:36 eileen: civicrm upgraded from {{Gerrit|1a835879}} to {{Gerrit|7556c5c7}} * 15:08 eileen: civicrm upgraded from {{Gerrit|9ed32632}} to {{Gerrit|1a835879}} * 14:02 eileen: civicrm upgraded from {{Gerrit|081d5a29}} to {{Gerrit|9ed32632}} == 2026-04-29 == * 13:51 jgleeson: payments-wiki upgraded from {{Gerrit|2e2eb8a2}} to {{Gerrit|4e0c944b}} * 13:49 jgleeson: tools upgraded from {{Gerrit|f52a5dcf}} to {{Gerrit|afbd0f67}} == 2026-04-28 == * 20:58 larssandergreen: civicrm upgraded from {{Gerrit|be3bb76b}} to {{Gerrit|081d5a29}} == 2026-04-27 == * 19:04 ejegg: fundraising civicrm upgraded from {{Gerrit|3f8d49fa}} to {{Gerrit|be3bb76b}} * 19:02 ejegg: payments-wiki upgraded from {{Gerrit|b1a352af}} to {{Gerrit|5265089d}} * 18:58 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|572b69da}} to {{Gerrit|88a1bcba}} == 2026-04-23 == * 20:40 ejegg: payments-wiki upgraded from {{Gerrit|6e78ef91}} to {{Gerrit|b1a352af}} * 19:52 ejegg: civicrm fundraising upgraded from {{Gerrit|5ea4c8d3}} to {{Gerrit|3f8d49fa}} * 19:27 ejegg: SmashPig upgraded from {{Gerrit|f1b3f3d9}} to {{Gerrit|572b69da}} * 16:32 larssandergreen: civicrm upgraded from {{Gerrit|53a0b46f}} to {{Gerrit|5ea4c8d3}} * 16:31 larssandergreen: tools upgraded from {{Gerrit|edca3f63}} to {{Gerrit|f52a5dcf}} == 2026-04-22 == * 02:10 eileen: civicrm upgraded from {{Gerrit|abd23ad7}} to {{Gerrit|53a0b46f}} == 2026-04-21 == * 23:16 cstone: civicrm upgraded from {{Gerrit|22f24ae4}} to {{Gerrit|abd23ad7}} * 19:46 larssandergreen: civicrm upgraded from {{Gerrit|ddc1f044}} to {{Gerrit|22f24ae4}} == 2026-04-20 == * 19:21 larssandergreen: tools upgraded from {{Gerrit|26ab0125}} to {{Gerrit|edca3f63}} * 15:09 ejegg: payments-wiki upgraded from {{Gerrit|86a42498}} to {{Gerrit|6e78ef91}} == 2026-04-17 == * 01:08 larssandergreen: civicrm upgraded from {{Gerrit|90c0ccd9}} to {{Gerrit|ddc1f044}} == 2026-04-16 == * 16:39 larssandergreen: tools upgraded from {{Gerrit|f14a814e}} to {{Gerrit|26ab0125}} * 14:20 larssandergreen: civicrm upgraded from {{Gerrit|801847a7}} to {{Gerrit|90c0ccd9}} * 14:19 larssandergreen: tools upgraded from {{Gerrit|9bff5f07}} to {{Gerrit|f14a814e}} * 02:39 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|61fee241}} to {{Gerrit|f1b3f3d9}} == 2026-04-15 == * 21:50 eileen: civicrm upgraded from {{Gerrit|6f33b6d0}} to {{Gerrit|801847a7}} * 06:14 eileen: civicrm upgraded from {{Gerrit|a047bf92}} to {{Gerrit|6f33b6d0}} * 01:20 eileen: SmashPig upgraded from {{Gerrit|100101fb}} to {{Gerrit|61fee241}} == 2026-04-14 == * 23:58 eileen: civicrm upgraded from {{Gerrit|eb3d73e4}} to {{Gerrit|a047bf92}} * 23:48 wfan: payments-wiki upgraded from {{Gerrit|26f5451a}} to {{Gerrit|c3b34f99}} * 22:31 eileen: civicrm upgraded from {{Gerrit|2058927e}} to {{Gerrit|eb3d73e4}} * 22:25 eileen: civicrm upgraded from {{Gerrit|2058927e}} to {{Gerrit|eb3d73e4}} * 18:06 ejegg: fundraising civicrm upgraded from {{Gerrit|fccf9b3a}} to {{Gerrit|2058927e}} == 2026-04-13 == * 18:20 ejegg: fundraising civicrm upgraded from {{Gerrit|fa20eb0a}} to {{Gerrit|fccf9b3a}} * 17:04 ejegg: re-enabled recurring donation charge jobs * 16:29 ejegg: fundraising civicrm upgraded from {{Gerrit|eb188fa2}} to {{Gerrit|fa20eb0a}} * 16:27 ejegg: disabled recurring donation charge jobs for code / settings update * 12:44 jgleeson: donorwiki upgraded from {{Gerrit|064a770e}} to {{Gerrit|26f5451a}} == 2026-04-10 == * 15:35 jgleeson: payments-wiki upgraded from {{Gerrit|c017d7e7}} to {{Gerrit|dd45f867}} == 2026-04-09 == * 23:31 wfan: payments-wiki upgraded from {{Gerrit|064a770e}} to {{Gerrit|c017d7e7}} * 21:49 ejegg: fundraising civicrm upgraded from {{Gerrit|3d3c0a62}} to {{Gerrit|eb188fa2}} * 19:00 larssandergreen: tools upgraded from {{Gerrit|986f7f83}} to {{Gerrit|9bff5f07}} * 13:08 jgleeson: civicrm upgraded from {{Gerrit|d8d3871c}} to {{Gerrit|3d3c0a62}} * 11:53 jgleeson: SmashPig upgraded from {{Gerrit|5c083891}} to {{Gerrit|100101fb}} * 01:20 ejegg: fundraising civicrm upgraded from {{Gerrit|e60321bb}} to {{Gerrit|d8d3871c}} == 2026-04-08 == * 17:44 ejegg: fundraising civicrm upgraded from {{Gerrit|4ee0b5e8}} to {{Gerrit|e60321bb}} * 15:12 ejegg: payments-wiki upgraded from {{Gerrit|1ad85e6c}} to {{Gerrit|064a770e}} * 01:19 ejegg: donorwiki upgraded from {{Gerrit|1ad85e6c}} to {{Gerrit|064a770e}} * 00:26 dwisehaupt: cloning new frdb frdb1008 from frdb2005 == 2026-04-07 == * 18:34 wfan: civicrm upgraded from {{Gerrit|9104e70b}} to {{Gerrit|6f762e29}} == 2026-04-06 == * 20:42 ejegg: re-enabled recurring donation charge job * 20:33 wfan: donorwiki upgraded from {{Gerrit|c2d03117}} to {{Gerrit|1ad85e6c}} * 20:32 wfan: payments-wiki upgraded from {{Gerrit|80cda166}} to {{Gerrit|1ad85e6c}} * 16:53 ejegg: disabled recurring donations charge job while diagnosing gr4vy routing errors * 16:03 ejegg: civicrm upgraded from {{Gerrit|4ee11209}} to {{Gerrit|9104e70b}} == 2026-04-03 == * 00:04 wfan: civicrm upgraded from {{Gerrit|49f541cd}} to {{Gerrit|4ee11209}} == 2026-04-02 == * 21:38 cstone: payments-wiki upgraded from {{Gerrit|86bec442}} to {{Gerrit|80cda166}} * 05:24 eileen: civicrm upgraded from {{Gerrit|c512abc6}} to {{Gerrit|49f541cd}} * 02:39 eileen: civicrm upgraded from {{Gerrit|bbed1291}} to {{Gerrit|c512abc6}} * 02:16 eileen: SmashPig upgraded from {{Gerrit|9af71a7c}} to {{Gerrit|18ea746a}} == 2026-04-01 == * 18:57 eileen: civicrm upgraded from {{Gerrit|a1bf4768}} to {{Gerrit|bbed1291}} * 04:11 eileen: civicrm upgraded from {{Gerrit|11a2f9ab}} to {{Gerrit|a1bf4768}} * 03:18 ejegg: payments-wiki upgraded from {{Gerrit|02bf54b0}} to {{Gerrit|86bec442}} == 2026-03-31 == * 22:03 jgleeson: tools upgraded from {{Gerrit|9985e723}} to {{Gerrit|986f7f83}} * 20:16 eileen: civicrm upgraded from {{Gerrit|c3cc3562}} to {{Gerrit|b468301c}} * 18:40 jgleeson: tools upgraded from {{Gerrit|161049ac}} to {{Gerrit|9985e723}} * 17:39 ejegg: Standalone (IPN listener) SmashPig upgraded from {{Gerrit|abf8682a}} to {{Gerrit|9af71a7c}} * 16:38 jgleeson: tools upgraded from {{Gerrit|f605b570}} to {{Gerrit|161049ac}} * 16:24 jgleeson: donorwiki updated from {{Gerrit|d79a98b5}} to {{Gerrit|c2d03117}} * 02:46 eileen: civicrm upgraded from {{Gerrit|591bef29}} to {{Gerrit|c3cc3562}} * 01:00 eileen: civicrm upgraded from {{Gerrit|cf871dd3}} to {{Gerrit|591bef29}} == 2026-03-30 == * 23:19 eileen: civicrm upgraded from {{Gerrit|7d299b48}} to {{Gerrit|cf871dd3}} * 21:09 eileen: civicrm upgraded from {{Gerrit|3724cc2d}} to {{Gerrit|7d299b48}} * 20:40 eileen: civicrm upgraded from {{Gerrit|58426b1e}} to {{Gerrit|3724cc2d}} * 14:51 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|545f0b10}} to {{Gerrit|abf8682a}} == 2026-03-28 == * 02:57 ejegg: payments-wiki upgraded from {{Gerrit|d79a98b5}} to {{Gerrit|b239a6b7}} * 02:27 eileen: civicrm upgraded from {{Gerrit|7138d524}} to {{Gerrit|58426b1e}} == 2026-03-27 == * 22:37 eileen: civicrm upgraded from {{Gerrit|c51e98cc}} to {{Gerrit|7138d524}} * 17:50 dwisehaupt: Corrected value: on frdb1005 running the following in mysql to up the buffer pool to ~412G: set global innodb_buffer_pool_size = 431669379072; * 17:34 dwisehaupt: on frdb1005 running the following in mysql to up the buffer pool to ~412G: set global innodb_buffer_pool_size = 421552128; == 2026-03-26 == * 20:00 jgleeson: donorwiki upgraded from {{Gerrit|48cc2e9d}} to {{Gerrit|d79a98b5}} * 19:59 jgleeson: payments-wiki upgraded from {{Gerrit|b387c6ba}} to {{Gerrit|d79a98b5}} * 19:48 ejegg: fundraising civicrm upgraded from {{Gerrit|88b497b8}} to {{Gerrit|db102c77}} * 19:25 larssandergreen: civicrm upgraded from {{Gerrit|26dab2f0}} to {{Gerrit|88b497b8}} * 16:26 jgleeson: SmashPig upgraded from {{Gerrit|5d8a0330}} to {{Gerrit|545f0b10}} * 02:13 eileen: civicrm upgraded from {{Gerrit|a752f9e8}} to {{Gerrit|26dab2f0}} * 00:05 eileen: civicrm upgraded from {{Gerrit|277fe75e}} to {{Gerrit|a752f9e8}} == 2026-03-25 == * 21:53 eileen: civicrm upgraded from {{Gerrit|97319295}} to {{Gerrit|277fe75e}} * 19:58 eileen: civicrm upgraded from {{Gerrit|d3dfedb4}} to {{Gerrit|97319295}} == 2026-03-24 == * 19:32 wfan: payments-wiki upgraded from {{Gerrit|fce3fab5}} to {{Gerrit|b387c6ba}} * 04:24 eileen: civicrm upgraded from {{Gerrit|0aef661f}} to {{Gerrit|b2ed875e}} * 03:38 eileen: config revision changed from {{Gerrit|ded0c289}} to {{Gerrit|16592428}} schedule stripe download * 03:33 eileen: config revision changed from {{Gerrit|79e052e4}} to {{Gerrit|ded0c289}} temporarily disable adyen audit parse - let's fix those misplaced IDs * 01:55 eileen: civicrm upgraded from {{Gerrit|b2c7f1d0}} to {{Gerrit|0aef661f}} * 00:08 eileen: civicrm upgraded from {{Gerrit|80344f51}} to {{Gerrit|b2c7f1d0}} == 2026-03-23 == * 21:30 eileen: config revision changed from {{Gerrit|8c5587f3}} to {{Gerrit|2dd50e7c}} * 18:55 wfan: civicrm upgraded from {{Gerrit|675455b2}} to {{Gerrit|80344f51}} * 17:27 larssandergreen: tools upgraded from {{Gerrit|e60f63b3}} to {{Gerrit|f605b570}} * 15:59 ejegg: civicrm upgraded from {{Gerrit|a2d4b17c}} to {{Gerrit|675455b2}} * 12:25 jgleeson: payments-wiki upgraded from {{Gerrit|48cc2e9d}} to {{Gerrit|91d9eee9}} == 2026-03-20 == * 00:34 eileen: * civicrm upgraded from {{Gerrit|adc36173}} to {{Gerrit|a2d4b17c}} * 00:31 cstone: payments-wiki upgraded from {{Gerrit|f3420a6f}} to {{Gerrit|48cc2e9d}} * 00:27 eileen: config revision changed from {{Gerrit|a7486f6a}} to {{Gerrit|a1a426f3}} * 00:27 eileen: SmashPig upgraded from {{Gerrit|78a8e70a}} to {{Gerrit|5d8a0330}} == 2026-03-19 == * 14:08 damilare: civiproxy upgraded from {{Gerrit|6625c844}} to {{Gerrit|38ba8348}} == 2026-03-17 == * 18:40 jgleeson: donorwiki upgraded from {{Gerrit|4c09db39}} to {{Gerrit|7d1666f9}} * 06:20 eileen: civicrm upgraded from {{Gerrit|7fe14629}} to {{Gerrit|adc36173}} * 05:06 eileen: civicrm upgraded from {{Gerrit|e622a222}} to {{Gerrit|7fe14629}} * 03:48 eileen: civicrm upgraded from {{Gerrit|5360f9ad}} to {{Gerrit|e622a222}} * 02:13 eileen: civicrm upgraded from {{Gerrit|3283e3ca}} to {{Gerrit|e73c6b50}} == 2026-03-15 == * 23:56 eileen: civicrm upgraded from {{Gerrit|dce257f0}} to {{Gerrit|3283e3ca}} * 19:43 eileen: civicrm upgraded from {{Gerrit|a1279ee4}} to {{Gerrit|dce257f0}} == 2026-03-13 == * ish: payments-wiki upgraded from {{Gerrit|f40a1153}} to {{Gerrit|f3420a6f}} == 2026-03-11 == * 21:20 larssandergreen: civicrm upgraded from {{Gerrit|c2c716ca}} to {{Gerrit|a1279ee4}} * 19:15 eileen: civicrm upgraded from {{Gerrit|81baf495}} to {{Gerrit|c2c716ca}} * 07:18 eileen: config revision changed from {{Gerrit|ed2295ab}} to {{Gerrit|a7486f6a}} * 07:02 eileen: civicrm upgraded from {{Gerrit|f418297f}} to {{Gerrit|81baf495}} * 02:28 eileen: civicrm upgraded from {{Gerrit|14e8200e}} to {{Gerrit|da26f37d}} * 00:06 eileen: civicrm upgraded from {{Gerrit|fbb38eda}} to {{Gerrit|14e8200e}} == 2026-03-10 == * 22:08 eileen: civicrm upgraded from {{Gerrit|ef319ea3}} to {{Gerrit|fbb38eda}} * 19:35 eileen: civicrm upgraded from {{Gerrit|773d9fb9}} to {{Gerrit|ef319ea3}} * 06:06 eileen: config revision changed from {{Gerrit|b9bc2a20}} to {{Gerrit|60ef6709}} == 2026-03-09 == * 17:37 damilare: payments-wiki upgraded from {{Gerrit|5b747b97}} to {{Gerrit|f40a1153}} == 2026-03-06 == * 23:39 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|ed16a2ea}} to {{Gerrit|78a8e70a}} * 21:48 ejegg: fundraising civicrm upgraded from {{Gerrit|8aadcd81}} to {{Gerrit|773d9fb9}} * 21:13 ejegg: civicrm fundraising upgraded from {{Gerrit|a1f32ed6}} to {{Gerrit|8aadcd81}} * 20:35 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|217fc7fc}} to {{Gerrit|ed16a2ea}} * 19:18 ejegg: civicrm upgraded from {{Gerrit|f2633c89}} to {{Gerrit|a1f32ed6}} * 03:08 larssandergreen: civicrm upgraded from {{Gerrit|fbac3ce7}} to {{Gerrit|2bae36fa}} * 02:09 ejegg: payments-wiki upgraded from {{Gerrit|9ae5bf60}} to {{Gerrit|5b747b97}} == 2026-03-05 == * 17:53 ejegg: donorwiki upgraded from {{Gerrit|7329b41d}} to {{Gerrit|4c09db39}} * 05:59 eileen: ivicrm upgraded from {{Gerrit|11e5a5d8}} to {{Gerrit|fbac3ce7}} * 05:08 eileen: * civicrm upgraded from {{Gerrit|8bdce85f}} to {{Gerrit|11e5a5d8}} == 2026-03-04 == * 18:52 jgleeson: tools upgraded from {{Gerrit|a3568ffc}} to {{Gerrit|e60f63b3}} * 03:21 ejegg: payments-wiki upgraded from {{Gerrit|5e4939a3}} to {{Gerrit|9ae5bf60}} * 03:20 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|78960c68}} to {{Gerrit|217fc7fc}} == 2026-03-03 == * 20:42 eileen: civicrm upgraded from {{Gerrit|b610f844}} to {{Gerrit|8bdce85f}} * 17:09 dwisehaupt: latest php8.2 updates installed on civi1002 * 06:00 eileen: * civicrm upgraded from {{Gerrit|f4a70c82}} to {{Gerrit|b610f844}} == 2026-03-02 == * 20:40 eileen: cv upgraded from {{Gerrit|dfeedcbe}} to {{Gerrit|f19e0961}} * 18:41 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|8cd1593b}} to {{Gerrit|78960c68}} * 15:13 ejegg: donorwiki upgraded from {{Gerrit|f5d7179a}} to {{Gerrit|7329b41d}} == 2026-02-27 == * 19:00 ejegg: fundraising civicrm upgraded from {{Gerrit|162bbf7c}} to {{Gerrit|f4a70c82}} * 12:24 jgleeson: payments-wiki upgraded from {{Gerrit|f4fd71ff}} to {{Gerrit|a8a9ce78}} * 11:25 jgleeson: payments-wiki upgraded from {{Gerrit|974af222}} to {{Gerrit|f4fd71ff}} == 2026-02-26 == * 22:36 eileen: civicrm upgraded from {{Gerrit|6995dc03}} to {{Gerrit|162bbf7c}} * 21:52 ejegg: civicrm upgraded from {{Gerrit|2a6001d3}} to {{Gerrit|6995dc03}} * 17:44 larssandergreen: civicrm upgraded from {{Gerrit|3ac93e95}} to {{Gerrit|2a6001d3}} * 07:14 eileen: civicrm upgraded from {{Gerrit|c3157fbf}} to {{Gerrit|3ac93e95}} == 2026-02-25 == * 18:33 ejegg: fundraising civicrm upgraded from {{Gerrit|f881e026}} to {{Gerrit|c3157fbf}} * 03:37 eileen: config revision changed from {{Gerrit|390d6434}} to {{Gerrit|a0228e6c}} turn off trustly audit == 2026-02-24 == * 22:48 ejegg: payments-wiki upgraded from {{Gerrit|e5f73610}} to {{Gerrit|974af222}} * 19:47 ejegg: payments-wiki upgraded from {{Gerrit|f5d7179a}} to {{Gerrit|e5f73610}} * 02:06 eileen: config revision changed from {{Gerrit|71c98072}} to {{Gerrit|390d6434}} reenabled trustly audit == 2026-02-23 == * 23:42 eileen: civicrm upgraded from {{Gerrit|f0710864}} to {{Gerrit|f881e026}} * 22:28 ejegg: fundraising civicrm upgraded from {{Gerrit|9d58ce4a}} to {{Gerrit|f0710864}} * 16:43 jgleeson: payments-wiki upgraded from {{Gerrit|0127f2d8}} to {{Gerrit|f5d7179a}} == 2026-02-21 == * 01:15 dwisehaupt: updating localsettings from {{Gerrit|71c98072}} to {{Gerrit|534fbf34}} and syncing civicrm to push update for large_donation_notifications == 2026-02-20 == * 20:38 ejegg: donorwiki upgraded from {{Gerrit|f7a0ee6b}} to {{Gerrit|f5d7179a}} == 2026-02-19 == * 21:26 ejegg: payments-wiki upgraded from {{Gerrit|f7a0ee6b}} to {{Gerrit|0127f2d8}} * 06:37 eileen: civicrm upgraded from {{Gerrit|ac30e19f}} to {{Gerrit|9d58ce4a}} == 2026-02-18 == * 22:21 dwisehaupt: disabling apache2::mod::dump_io on civicrm role for debugging 500 errors after testing. can be re-enabled by reverting commit {{Gerrit|4b1d94399}} - [[phab:T417310|T417310]] * 20:38 eileen: civicrm upgraded from {{Gerrit|f5020a85}} to {{Gerrit|ac30e19f}} * 19:22 eileen: civicrm upgraded from {{Gerrit|66d2e1dd}} to {{Gerrit|f5020a85}} * 17:48 ejegg: donorwiki upgraded from {{Gerrit|488431ec}} to {{Gerrit|f7a0ee6b}} * 16:51 dwisehaupt: enabling apache2::mod::dump_io on civicrm role for debugging 500 errors - [[phab:T417310|T417310]] * 06:02 eileen: civicrm upgraded from {{Gerrit|4c3cdcde}} to {{Gerrit|66d2e1dd}} * 05:33 eileen: config revision changed from {{Gerrit|605c6946}} to {{Gerrit|368156fa}} * 05:21 eileen: civicrm upgraded from {{Gerrit|caad5ab9}} to {{Gerrit|4c3cdcde}} == 2026-02-17 == * 17:34 larssandergreen: payments-wiki upgraded from {{Gerrit|c506d590}} to {{Gerrit|488431ec}} * 17:33 larssandergreen: donorwiki upgraded from {{Gerrit|93e0d03f}} to {{Gerrit|488431ec}} * 03:02 eileen: civicrm upgraded from {{Gerrit|89782fc6}} to {{Gerrit|caad5ab9}} == 2026-02-16 == * 20:16 eileen: civicrm upgraded from {{Gerrit|2b227403}} to {{Gerrit|89782fc6}} == 2026-02-15 == * 23:55 eileen: civicrm upgraded from {{Gerrit|de8252c7}} to {{Gerrit|2b227403}} * 21:07 eileen: civicrm upgraded from {{Gerrit|038f5bca}} to {{Gerrit|de8252c7}} == 2026-02-13 == * 16:03 ejegg: payments-wiki upgraded from {{Gerrit|5793a405}} to {{Gerrit|c506d590}} * away: SmashPig upgraded from {{Gerrit|fea03fcc}} to {{Gerrit|8cd1593b}} == 2026-02-12 == * 22:49 eileen: civicrm upgraded from {{Gerrit|c6c0d453}} to {{Gerrit|038f5bca}} * 21:45 jgleeson: payments-wiki upgraded from {{Gerrit|9dbf0ece}} to {{Gerrit|5793a405}} * 21:15 jgleeson: payments-wiki upgraded from {{Gerrit|6c1a522f}} to {{Gerrit|9dbf0ece}} * 19:03 larssandergreen: tools upgraded from {{Gerrit|645cf5dc}} to {{Gerrit|a3568ffc}} * 00:34 larssandergreen: civicrm upgraded from {{Gerrit|e13111f3}} to {{Gerrit|c6c0d453}} == 2026-02-11 == * 21:12 eileen: civicrm upgraded from {{Gerrit|6e57071a}} to {{Gerrit|e13111f3}} * 06:27 eileen: civicrm upgraded from {{Gerrit|98c325dd}} to {{Gerrit|6e57071a}} * 04:55 larssandergreen: tools upgraded from {{Gerrit|7462b8bd}} to {{Gerrit|645cf5dc}} * 03:37 eileen: civicrm upgraded from {{Gerrit|953cf9f2}} to {{Gerrit|98c325dd}} == 2026-02-09 == * 15:35 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|937d6e40}} to {{Gerrit|fea03fcc}} == 2026-02-08 == * 23:47 eileen: civicrm upgraded from {{Gerrit|bc3a8036}} to {{Gerrit|953cf9f2}} * 20:46 eileen: civicrm upgraded from {{Gerrit|40189afa}} to {{Gerrit|bc3a8036}} == 2026-02-06 == * 15:41 ejegg: fundraising civicrm upgraded from {{Gerrit|06842fbf}} to {{Gerrit|40189afa}} * 06:24 eileen: civicrm upgraded from {{Gerrit|b63d7146}} to {{Gerrit|06842fbf}} == 2026-02-05 == * 15:25 dwisehaupt: all eqiad hosts powered down and ready for relocation. * 15:03 dwisehaupt: starting poweroff of eqiad hosts * 14:48 dwisehaupt: downtimes scheduled for frack eqiad hosts and cross colo replication in prep for rack expansion - [[phab:T403035|T403035]] * 04:30 eileen: civicrm upgraded from {{Gerrit|000dd548}} to {{Gerrit|b63d7146}} == 2026-02-04 == * 23:56 eileen: civicrm upgraded from {{Gerrit|4c2870c5}} to {{Gerrit|000dd548}} * 23:47 cstone: donorwiki upgraded from {{Gerrit|53bfb05b}} to {{Gerrit|93e0d03f}} * 21:51 ejegg: payments-wiki upgraded from {{Gerrit|a09a4f8f}} to {{Gerrit|93e0d03f}} * 21:46 eileen: civicrm upgraded from {{Gerrit|dd10342d}} to {{Gerrit|4c2870c5}} * 21:22 eileen: civicrm upgraded from {{Gerrit|8aa9274a}} to {{Gerrit|dd10342d}} * 03:48 eileen: * civicrm upgraded from {{Gerrit|14c7b7e7}} to {{Gerrit|8aa9274a}} == 2026-02-03 == * 21:55 dwisehaupt: as part of dns cleanup, we have removed the old civi1002.wikimedia.org entry. folks should be using civicrm.wm.o but there is a chance of super old bookmarks still being around. we can reinstate if it becomes an issue. * 21:05 eileen: config revision changed from {{Gerrit|23b2d9b6}} to {{Gerrit|45d40cf1}} * 21:01 larssandergreen: civicrm upgraded from {{Gerrit|10ab3659}} to {{Gerrit|14c7b7e7}} * 05:49 eileen: config revision changed from {{Gerrit|f348441d}} to {{Gerrit|23b2d9b6}} * 05:48 eileen: civicrm upgraded from {{Gerrit|a097bb3d}} to {{Gerrit|10ab3659}} * 01:16 wfan: civicrm upgraded from {{Gerrit|5b9f3cd4}} to {{Gerrit|a097bb3d}} * 00:24 wfan: donorwiki upgraded from {{Gerrit|3ffc70f0}} to {{Gerrit|53bfb05b}} == 2026-02-02 == * 23:50 wfan: payments-wiki upgraded from {{Gerrit|c035aa84}} to {{Gerrit|53bfb05b}} * 20:22 cstone: civicrm upgraded from {{Gerrit|f91f955b}} to {{Gerrit|5b9f3cd4}} * 18:55 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|b42f2de4}} to {{Gerrit|937d6e40}} * 17:01 wfan: payments-wiki upgraded from {{Gerrit|c5cadd72}} to {{Gerrit|c035aa84}} * 15:26 ejegg: fundraising civicrm upgraded from {{Gerrit|611d18de}} to {{Gerrit|f91f955b}} == 2026-01-30 == * 17:02 larssandergreen: civicrm upgraded from {{Gerrit|79e4424e}} to {{Gerrit|611d18de}} * 04:18 cstone: civicrm upgraded from {{Gerrit|fe1af57a}} to {{Gerrit|79e4424e}} * 02:51 eileen: civicrm upgraded from {{Gerrit|ff772bee}} to {{Gerrit|fe1af57a}} * 02:28 eileen: civicrm upgraded from {{Gerrit|bcf976ae}} to {{Gerrit|ff772bee}} == 2026-01-29 == * 22:16 eileen: civicrm upgraded from {{Gerrit|5d121c63}} to {{Gerrit|bcf976ae}} * 20:17 eileen: civicrm upgraded from {{Gerrit|ebcfd009}} to {{Gerrit|5d121c63}} * 20:09 eileen: civicrm upgraded from {{Gerrit|5c065c4e}} to {{Gerrit|ebcfd009}} * 18:08 wfan: payments-wiki upgraded from {{Gerrit|81d9f614}} to {{Gerrit|c5cadd72}} == 2026-01-28 == * 19:36 jgleeson: SmashPig upgraded from {{Gerrit|96a6224d}} to {{Gerrit|b42f2de4}} * 18:53 jgleeson: civicrm upgraded from {{Gerrit|56c222da}} to {{Gerrit|5c065c4e}} * 18:53 jgleeson: SmashPig upgraded from {{Gerrit|96a6224d}} to {{Gerrit|b42f2de4}} * 14:41 jgleeson: payments-wiki upgraded from {{Gerrit|24915bdb}} to {{Gerrit|81d9f614}} * 07:30 eileen: civicrm upgraded from {{Gerrit|600b21a6}} to {{Gerrit|56c222da}} * 06:51 eileen: * civicrm upgraded from {{Gerrit|32f9a10d}} to {{Gerrit|600b21a6}} * 04:35 eileen: config revision changed from {{Gerrit|ed0808a9}} to {{Gerrit|ef6ef5f2}} * 03:42 eileen: civicrm upgraded from {{Gerrit|64267a34}} to {{Gerrit|32f9a10d}} * 02:13 eileen: civicrm upgraded from {{Gerrit|7299615a}} to {{Gerrit|64267a34}} == 2026-01-27 == * 21:12 larssandergreen: tools upgraded from {{Gerrit|84323460}} to {{Gerrit|7462b8bd}} * 05:57 cstone: civicrm upgraded from {{Gerrit|19f94835}} to {{Gerrit|75f443b5}} == 2026-01-26 == * 17:15 damilare: smashpig upgraded from {{Gerrit|8b4ebf34}} to {{Gerrit|96a6224d}} * 16:32 larssandergreen: tools upgraded from {{Gerrit|c75f7625}} to {{Gerrit|84323460}} * 01:43 eileen: config revision changed from {{Gerrit|2f71107f}} to {{Gerrit|ed0808a9}} switch to php dlocal downloader (now weekend is mostly over) - == 2026-01-25 == * 21:47 eileen: config revision changed from {{Gerrit|23023984}} to {{Gerrit|2f71107f}} == 2026-01-24 == * 02:19 cstone: civicrm upgraded from {{Gerrit|f7064a46}} to {{Gerrit|19f94835}} * 00:55 bd808: Testing #wikimedia-fundraising SAL integration ([[phab:T415389|T415389]]) <noinclude>[[Category:SAL]]</noinclude> cktl5lxwi3l2l2cfe7p4qtb2p8h7hmm 2445269 2445268 2026-08-10T04:23:05Z Stashbot 7414 eileen: civicrm revision 78e2ecac -> 0947b202 - about 15 mins ago but didn't log at the right time 2445269 wikitext text/x-wiki == 2026-08-10 == * 04:23 eileen: civicrm revision {{Gerrit|78e2ecac}} -> {{Gerrit|0947b202}} - about 15 mins ago but didn't log at the right time * 04:21 eileen: config revision changed from {{Gerrit|e99fd679}} to {{Gerrit|4b91cb74}} == 2026-08-06 == * 14:44 larssandergreen: civicrm upgraded from {{Gerrit|56c8de1f}} to {{Gerrit|78e2ecac}} * 01:36 larssandergreen: civicrm upgraded from {{Gerrit|6408e93a}} to {{Gerrit|56c8de1f}} == 2026-08-05 == * 17:53 larssandergreen: civicrm upgraded from {{Gerrit|e229c253}} to {{Gerrit|6408e93a}} == 2026-08-04 == * 23:15 eileen: civicrm upgraded from {{Gerrit|8ed3282e}} to {{Gerrit|e229c253}} * 02:35 ejegg: fundraising civicrm upgraded from {{Gerrit|9c8ee02f}} to {{Gerrit|8ed3282e}} * 00:51 eileen: checking attributes == 2026-08-03 == * 22:27 eileen: civicrm upgraded from {{Gerrit|26bd5ab6}} to {{Gerrit|9c8ee02f}} * 22:06 eileen: SmashPig upgraded from {{Gerrit|0c2593b5}} to {{Gerrit|c89520d4}} == 2026-07-30 == * 13:53 damilare: donorwiki upgraded from {{Gerrit|d59684ed}} to {{Gerrit|0117d21d}} == 2026-07-28 == * 20:08 ejegg: fundraising civicrm upgraded from {{Gerrit|98ef4abd}} to {{Gerrit|26bd5ab6}} * 16:34 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|038018d4}} to {{Gerrit|0c2593b5}} == 2026-07-23 == * 23:50 larssandergreen: civicrm upgraded from {{Gerrit|bf93abac}} to {{Gerrit|98ef4abd}} * 00:04 eileen: civicrm upgraded from {{Gerrit|068b2e3c}} to {{Gerrit|bf93abac}} == 2026-07-22 == * 22:35 eileen: civicrm upgraded from {{Gerrit|5a04e8a9}} to {{Gerrit|068b2e3c}} * 21:38 eileen: SmashPig upgraded from {{Gerrit|6921d862}} to {{Gerrit|038018d4}} * 21:36 eileen: SmashPig upgraded from {{Gerrit|6921d862}} to {{Gerrit|038018d4}} * 01:36 eileen: civicrm upgraded from {{Gerrit|2eba4374}} to {{Gerrit|5a04e8a9}} == 2026-07-17 == * 03:44 eileen: civicrm upgraded from {{Gerrit|dd65e1b6}} to {{Gerrit|2eba4374}} * 01:25 eileen: civicrm upgraded from {{Gerrit|1c94b489}} to {{Gerrit|dd65e1b6}} == 2026-07-16 == * 11:06 laurabar: process-control config revision changed from {{Gerrit|730715e1}} to {{Gerrit|c951e730}} * 10:41 laurabar: process-control config revision changed from {{Gerrit|7dddea38}} to {{Gerrit|730715e1}} * 02:31 eileen: civicrm upgraded from {{Gerrit|2dd23034}} to {{Gerrit|1c94b489}} * 01:51 eileen: SmashPig upgraded from {{Gerrit|2e07ce09}} to {{Gerrit|6921d862}} == 2026-07-15 == * 23:48 eileen: config revision changed from {{Gerrit|8078b0b0}} to {{Gerrit|7dddea38}} * 23:12 eileen: SmashPig upgraded from {{Gerrit|37a80ec4}} to {{Gerrit|2e07ce09}} * 21:48 eileen: civicrm upgraded from {{Gerrit|d38ae4d6}} to {{Gerrit|2dd23034}} * 21:19 eileen: SmashPig upgraded from {{Gerrit|34082cca}} to {{Gerrit|37a80ec4}} * 04:58 eileen: civicrm upgraded from {{Gerrit|24d6e9a0}} to {{Gerrit|d38ae4d6}} * 04:30 eileen: SmashPig upgraded from {{Gerrit|87dfbd6b}} to {{Gerrit|34082cca}} * 03:44 cstone: payments-wiki upgraded from {{Gerrit|776f7d37}} to {{Gerrit|5482106f}} == 2026-07-14 == * 02:45 larssandergreen: tools upgraded from {{Gerrit|0c0cbb37}} to {{Gerrit|93a9d78d}} * 01:42 larssandergreen: civicrm upgraded from {{Gerrit|015e84b8}} to {{Gerrit|24d6e9a0}} == 2026-07-13 == * 22:46 larssandergreen: civicrm upgraded from {{Gerrit|171652fa}} to {{Gerrit|015e84b8}} * 19:48 eileen: civicrm upgraded from {{Gerrit|263ff140}} to {{Gerrit|171652fa}} == 2026-07-10 == * 21:17 larssandergreen: tools upgraded from {{Gerrit|afbd0f67}} to {{Gerrit|0c0cbb37}} * 20:00 cstone: civicrm upgraded from {{Gerrit|20b3dabe}} to {{Gerrit|263ff140}} == 2026-07-09 == * 05:02 eileen: civicrm upgraded from {{Gerrit|3a1a0891}} to {{Gerrit|20b3dabe}} * 03:12 eileen: civicrm upgraded from {{Gerrit|3a2952ff}} to {{Gerrit|3a1a0891}} * 02:28 eileen: civicrm upgraded from {{Gerrit|f4ff79c8}} to {{Gerrit|3a2952ff}} * 02:12 eileen: SmashPig upgraded from {{Gerrit|2042f7f7}} to {{Gerrit|87dfbd6b}} * 00:01 eileen: onfig {{Gerrit|1c59ae42}} -> {{Gerrit|fa79267c}} == 2026-07-08 == * 23:52 eileen: SmashPig upgraded from {{Gerrit|d8393666}} to {{Gerrit|2042f7f7}} * 22:15 larssandergreen: civicrm upgraded from {{Gerrit|ead61459}} to {{Gerrit|f4ff79c8}} * 22:13 larssandergreen: donorwiki upgraded from {{Gerrit|c2cdbf24}} to {{Gerrit|d59684ed}} * 19:50 jgleeson: payments-wiki upgraded from {{Gerrit|196c7bc1}} to {{Gerrit|5d2bb8e8}} * 13:46 jgleeson: payments-wiki upgraded from {{Gerrit|f34901d7}} to {{Gerrit|196c7bc1}} * 03:48 larssandergreen: civicrm upgraded from {{Gerrit|0057f9a1}} to {{Gerrit|ead61459}} * 01:12 larssandergreen: civicrm upgraded from {{Gerrit|941505c3}} to {{Gerrit|0057f9a1}} == 2026-07-07 == * 22:29 jgleeson: payments-wiki upgraded from {{Gerrit|ece97dca}} to {{Gerrit|f34901d7}} * 04:47 eileen: civicrm upgraded from {{Gerrit|0013ea5e}} to {{Gerrit|941505c3}} * 03:27 eileen: process-control * 01:35 larssandergreen: civicrm upgraded from {{Gerrit|274308a4}} to {{Gerrit|0013ea5e}} == 2026-07-02 == * 23:46 eileen: civicrm upgraded from {{Gerrit|d44caa2a}} to {{Gerrit|274308a4}} * 03:36 eileen: civicrm upgraded from {{Gerrit|7288da0a}} to {{Gerrit|d44caa2a}} * 01:54 eileen: * civicrm upgraded from {{Gerrit|13584dc4}} to {{Gerrit|7288da0a}} * 01:46 eileen: config revision changed from {{Gerrit|92ac127c}} to {{Gerrit|d0a8b49f}} == 2026-06-30 == * 21:25 eileen: SmashPig upgraded from {{Gerrit|38ffd696}} to {{Gerrit|d8393666}} * 07:01 eileen: civicrm upgraded from {{Gerrit|63606ee8}} to {{Gerrit|c2fa32a6}} == 2026-06-29 == * 23:16 danielfm: payments-wiki upgraded from {{Gerrit|7e1a422e}} to {{Gerrit|28983fa4}} * 18:54 ejegg: payments-wiki upgraded from {{Gerrit|32059801}} to {{Gerrit|7e1a422e}} * 18:32 wfan: payments-wiki upgraded from {{Gerrit|39df7f6b}} to {{Gerrit|32059801}} * 18:08 ejegg: payments-wiki upgraded from {{Gerrit|c2cdbf24}} to {{Gerrit|39df7f6b}} == 2026-06-26 == * 08:17 eileen: civicrm upgraded from {{Gerrit|a90a5451}} to {{Gerrit|ebab84c8}} == 2026-06-25 == * 13:35 jgleeson: payments-wiki upgraded from {{Gerrit|ab4fd16b}} to {{Gerrit|c2cdbf24}} == 2026-06-24 == * 17:30 ejegg: fundraising civicrm upgraded from {{Gerrit|7130d7ff}} to {{Gerrit|a90a5451}} == 2026-06-23 == * 09:29 jgleeson: payments-wiki upgraded from {{Gerrit|71cba440}} to {{Gerrit|ab4fd16b}} == 2026-06-22 == * 22:53 eileen: SmashPig upgraded from {{Gerrit|9cd51fd1}} to {{Gerrit|38ffd696}} * 22:32 eileen: config revision changed from {{Gerrit|c687d8f0}} to {{Gerrit|92ac127c}} * 22:15 eileen: civicrm upgraded from {{Gerrit|602742db}} to {{Gerrit|7130d7ff}} * 20:27 ejegg: donorwiki upgraded from {{Gerrit|1f5f40f9}} to {{Gerrit|71cba440}} * 16:35 ejegg: payments-wiki upgraded from {{Gerrit|873882a5}} to {{Gerrit|d5f4b8aa}} == 2026-06-18 == * 00:42 eileen: civicrm upgraded from {{Gerrit|646d893e}} to {{Gerrit|602742db}} == 2026-06-17 == * 01:22 ejegg: payments-wiki upgraded from {{Gerrit|1f5f40f9}} to {{Gerrit|873882a5}} == 2026-06-16 == * 20:31 ejegg: standalone (ipn listener) SmashPig upgraded from {{Gerrit|611947be}} to {{Gerrit|9cd51fd1}} * 03:02 eileen: SmashPig upgraded from {{Gerrit|8ad963c7}} to {{Gerrit|611947be}} * 01:16 eileen: config revision changed from {{Gerrit|7ca9f992}} to {{Gerrit|8334c030}} * 01:12 eileen: updated SmashPig revision {{Gerrit|e82c2c5f}} -> {{Gerrit|8ad963c7}} == 2026-06-15 == * 11:43 damilare: donorwiki upgraded from {{Gerrit|3bc70a73}} to {{Gerrit|1f5f40f9}} == 2026-06-12 == * 18:08 jgleeson: civicrm upgraded from {{Gerrit|69a60dcb}} to {{Gerrit|646d893e}} * 03:42 eileen: civicrm upgraded from {{Gerrit|0b8db587}} to {{Gerrit|69a60dcb}} == 2026-06-11 == * 18:56 jgleeson: payments-wiki upgraded from {{Gerrit|aef3d25d}} to {{Gerrit|1f5f40f9}} * 06:05 eileen: civicrm upgraded from {{Gerrit|77961b3e}} to {{Gerrit|0b8db587}} * 02:59 larssandergreen: civicrm upgraded from {{Gerrit|8f98770f}} to {{Gerrit|77961b3e}} * 02:42 eileen: civicrm upgraded from {{Gerrit|2962b28d}} to {{Gerrit|8f98770f}} * 01:30 eileen: civicrm upgraded from {{Gerrit|f46a6066}} to {{Gerrit|2962b28d}} * 00:22 eileen: civicrm upgraded from {{Gerrit|819c4ede}} to {{Gerrit|f46a6066}} == 2026-06-10 == * 16:04 wfan: civicrm upgraded from {{Gerrit|d7113e04}} to {{Gerrit|819c4ede}} * 11:36 jgleeson: SmashPig upgraded from {{Gerrit|f6b24f7f}} to {{Gerrit|e82c2c5f}} == 2026-06-09 == * 21:39 eileen: civicrm upgraded from {{Gerrit|6b61d3a5}} to {{Gerrit|d7113e04}} == 2026-06-04 == * 11:32 jgleeson: payments-wiki upgraded from {{Gerrit|3bc70a73}} to {{Gerrit|aef3d25d}} == 2026-06-03 == * 12:17 jgleeson: SmashPig upgraded from {{Gerrit|166abfbd}} to {{Gerrit|99233b18}} * 05:46 eileen: civicrm upgraded from {{Gerrit|55adc0bb}} to {{Gerrit|6b61d3a5}} * 04:28 eileen: civicrm upgraded from {{Gerrit|219cf085}} to {{Gerrit|55adc0bb}} * 03:32 eileen: civicrm upgraded from {{Gerrit|f6839b65}} to {{Gerrit|219cf085}} * 02:07 eileen: civicrm upgraded from {{Gerrit|663c0c30}} to {{Gerrit|f6839b65}} == 2026-06-02 == * 23:14 eileen: config revision changed from {{Gerrit|67231927}} to {{Gerrit|c011a0d6}} * 22:55 eileen: config revision changed from {{Gerrit|46874905}} to {{Gerrit|67231927}} == 2026-06-01 == * 19:53 eileen: civicrm upgraded from {{Gerrit|a4a885f4}} to {{Gerrit|663c0c30}} * 13:20 damilare: donorwiki upgraded from {{Gerrit|9f3734e0}} to {{Gerrit|3bc70a73}} == 2026-05-28 == * 22:09 eileen: SmashPig upgraded from {{Gerrit|252b6cfa}} to {{Gerrit|166abfbd}} * 19:44 dwisehaupt: removing apache2::mod::python from civicrm and frdev roles. * 18:43 ejegg: payments-wiki upgraded from {{Gerrit|9f3734e0}} to {{Gerrit|21185532}} * 15:13 ejegg: donorwiki upgraded from {{Gerrit|1a056dc0}} to {{Gerrit|9f3734e0}} * 15:12 ejegg: payments-wiki upgraded from {{Gerrit|1e2eb148}} to {{Gerrit|9f3734e0}} * 14:38 ejegg: civicrm upgraded from {{Gerrit|0f0567b3}} to {{Gerrit|a4a885f4}} == 2026-05-26 == * 01:05 eileen: civicrm upgraded from {{Gerrit|6b4e53a5}} to {{Gerrit|713e8508}} == 2026-05-21 == * 04:13 eileen: civicrm upgraded from {{Gerrit|cae2ab0e}} to {{Gerrit|dbafc0b4}} * 00:09 cstone: civicrm upgraded from {{Gerrit|fcbbf763}} to {{Gerrit|cae2ab0e}} == 2026-05-20 == * 22:22 eileen: SmashPig upgraded from {{Gerrit|df441f78}} to {{Gerrit|252b6cfa}} == 2026-05-19 == * 23:08 eileen: civicrm upgraded from {{Gerrit|6cff86c1}} to {{Gerrit|fcbbf763}} * 01:56 eileen: civicrm upgraded from {{Gerrit|bfa13f7f}} to {{Gerrit|6cff86c1}} == 2026-05-18 == * 20:45 eileen: civicrm upgraded from {{Gerrit|150d484e}} to {{Gerrit|bfa13f7f}} * 14:26 ejegg: payments-wiki upgraded from {{Gerrit|40a24102}} to {{Gerrit|e31ac2a9}} * 11:29 jgleeson: payments-wiki upgraded from {{Gerrit|1a056dc0}} to {{Gerrit|40a24102}} == 2026-05-15 == * 19:19 dwisehaupt: redis swap complete from frqueue1003 to frqueue1005 * 15:59 ejegg: donorwiki upgraded from {{Gerrit|26f5451a}} to {{Gerrit|1a056dc0}} * 15:35 ejegg: payments-wiki upgraded from {{Gerrit|cf9ec80b}} to {{Gerrit|1a056dc0}} * 00:02 eileen: civicrm upgraded from {{Gerrit|6d8ce7a3}} to {{Gerrit|6a2258ff}} == 2026-05-14 == * 22:22 ejegg: fundraising scheduled jobs re-enabled * 22:12 eileen: cv upgraded from {{Gerrit|f19e0961}} to {{Gerrit|b8a8dd6a}} * 22:11 ejegg: fundraising civicrm upgraded from {{Gerrit|e25fa223}} to {{Gerrit|6d8ce7a3}} * 22:09 ejegg: fundraising scheduled jobs disabled for Civi update * 19:29 ejegg: re-enabled fundraising scheduled jobs * 19:06 ejegg: fundraising civicrm upgraded from {{Gerrit|950908ec}} to {{Gerrit|e25fa223}} * 19:04 ejegg: disabled fundraising scheduled jobs for CiviCRM deployment * 16:48 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|bb833986}} to {{Gerrit|a2a3a015}} == 2026-05-12 == * 16:07 ejegg: fundraising civicrm upgraded from {{Gerrit|24ac90e7}} to {{Gerrit|60dcc28f}} == 2026-05-06 == * 12:29 eileen: config revision changed from {{Gerrit|41cfd677}} to {{Gerrit|00752f91}} * 10:56 eileen: SmashPig upgraded from {{Gerrit|4201ef56}} to {{Gerrit|bb833986}} * 09:59 eileen: civicrm upgraded from {{Gerrit|4d9c8600}} to {{Gerrit|24ac90e7}} * 09:15 eileen: civicrm upgraded from {{Gerrit|38dcf7a8}} to {{Gerrit|4d9c8600}} == 2026-05-04 == * 17:39 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|1be60746}} to {{Gerrit|4201ef56}} == 2026-05-03 == * hackathon: civicrm upgraded from {{Gerrit|0afbc8ea}} to {{Gerrit|38dcf7a8}} == 2026-05-02 == * 09:44 eileen: civicrm upgraded from {{Gerrit|7556c5c7}} to {{Gerrit|0afbc8ea}} * 09:43 eileen: SmashPig upgraded from {{Gerrit|88a1bcba}} to {{Gerrit|1be60746}} == 2026-05-01 == * 16:36 eileen: civicrm upgraded from {{Gerrit|1a835879}} to {{Gerrit|7556c5c7}} * 15:08 eileen: civicrm upgraded from {{Gerrit|9ed32632}} to {{Gerrit|1a835879}} * 14:02 eileen: civicrm upgraded from {{Gerrit|081d5a29}} to {{Gerrit|9ed32632}} == 2026-04-29 == * 13:51 jgleeson: payments-wiki upgraded from {{Gerrit|2e2eb8a2}} to {{Gerrit|4e0c944b}} * 13:49 jgleeson: tools upgraded from {{Gerrit|f52a5dcf}} to {{Gerrit|afbd0f67}} == 2026-04-28 == * 20:58 larssandergreen: civicrm upgraded from {{Gerrit|be3bb76b}} to {{Gerrit|081d5a29}} == 2026-04-27 == * 19:04 ejegg: fundraising civicrm upgraded from {{Gerrit|3f8d49fa}} to {{Gerrit|be3bb76b}} * 19:02 ejegg: payments-wiki upgraded from {{Gerrit|b1a352af}} to {{Gerrit|5265089d}} * 18:58 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|572b69da}} to {{Gerrit|88a1bcba}} == 2026-04-23 == * 20:40 ejegg: payments-wiki upgraded from {{Gerrit|6e78ef91}} to {{Gerrit|b1a352af}} * 19:52 ejegg: civicrm fundraising upgraded from {{Gerrit|5ea4c8d3}} to {{Gerrit|3f8d49fa}} * 19:27 ejegg: SmashPig upgraded from {{Gerrit|f1b3f3d9}} to {{Gerrit|572b69da}} * 16:32 larssandergreen: civicrm upgraded from {{Gerrit|53a0b46f}} to {{Gerrit|5ea4c8d3}} * 16:31 larssandergreen: tools upgraded from {{Gerrit|edca3f63}} to {{Gerrit|f52a5dcf}} == 2026-04-22 == * 02:10 eileen: civicrm upgraded from {{Gerrit|abd23ad7}} to {{Gerrit|53a0b46f}} == 2026-04-21 == * 23:16 cstone: civicrm upgraded from {{Gerrit|22f24ae4}} to {{Gerrit|abd23ad7}} * 19:46 larssandergreen: civicrm upgraded from {{Gerrit|ddc1f044}} to {{Gerrit|22f24ae4}} == 2026-04-20 == * 19:21 larssandergreen: tools upgraded from {{Gerrit|26ab0125}} to {{Gerrit|edca3f63}} * 15:09 ejegg: payments-wiki upgraded from {{Gerrit|86a42498}} to {{Gerrit|6e78ef91}} == 2026-04-17 == * 01:08 larssandergreen: civicrm upgraded from {{Gerrit|90c0ccd9}} to {{Gerrit|ddc1f044}} == 2026-04-16 == * 16:39 larssandergreen: tools upgraded from {{Gerrit|f14a814e}} to {{Gerrit|26ab0125}} * 14:20 larssandergreen: civicrm upgraded from {{Gerrit|801847a7}} to {{Gerrit|90c0ccd9}} * 14:19 larssandergreen: tools upgraded from {{Gerrit|9bff5f07}} to {{Gerrit|f14a814e}} * 02:39 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|61fee241}} to {{Gerrit|f1b3f3d9}} == 2026-04-15 == * 21:50 eileen: civicrm upgraded from {{Gerrit|6f33b6d0}} to {{Gerrit|801847a7}} * 06:14 eileen: civicrm upgraded from {{Gerrit|a047bf92}} to {{Gerrit|6f33b6d0}} * 01:20 eileen: SmashPig upgraded from {{Gerrit|100101fb}} to {{Gerrit|61fee241}} == 2026-04-14 == * 23:58 eileen: civicrm upgraded from {{Gerrit|eb3d73e4}} to {{Gerrit|a047bf92}} * 23:48 wfan: payments-wiki upgraded from {{Gerrit|26f5451a}} to {{Gerrit|c3b34f99}} * 22:31 eileen: civicrm upgraded from {{Gerrit|2058927e}} to {{Gerrit|eb3d73e4}} * 22:25 eileen: civicrm upgraded from {{Gerrit|2058927e}} to {{Gerrit|eb3d73e4}} * 18:06 ejegg: fundraising civicrm upgraded from {{Gerrit|fccf9b3a}} to {{Gerrit|2058927e}} == 2026-04-13 == * 18:20 ejegg: fundraising civicrm upgraded from {{Gerrit|fa20eb0a}} to {{Gerrit|fccf9b3a}} * 17:04 ejegg: re-enabled recurring donation charge jobs * 16:29 ejegg: fundraising civicrm upgraded from {{Gerrit|eb188fa2}} to {{Gerrit|fa20eb0a}} * 16:27 ejegg: disabled recurring donation charge jobs for code / settings update * 12:44 jgleeson: donorwiki upgraded from {{Gerrit|064a770e}} to {{Gerrit|26f5451a}} == 2026-04-10 == * 15:35 jgleeson: payments-wiki upgraded from {{Gerrit|c017d7e7}} to {{Gerrit|dd45f867}} == 2026-04-09 == * 23:31 wfan: payments-wiki upgraded from {{Gerrit|064a770e}} to {{Gerrit|c017d7e7}} * 21:49 ejegg: fundraising civicrm upgraded from {{Gerrit|3d3c0a62}} to {{Gerrit|eb188fa2}} * 19:00 larssandergreen: tools upgraded from {{Gerrit|986f7f83}} to {{Gerrit|9bff5f07}} * 13:08 jgleeson: civicrm upgraded from {{Gerrit|d8d3871c}} to {{Gerrit|3d3c0a62}} * 11:53 jgleeson: SmashPig upgraded from {{Gerrit|5c083891}} to {{Gerrit|100101fb}} * 01:20 ejegg: fundraising civicrm upgraded from {{Gerrit|e60321bb}} to {{Gerrit|d8d3871c}} == 2026-04-08 == * 17:44 ejegg: fundraising civicrm upgraded from {{Gerrit|4ee0b5e8}} to {{Gerrit|e60321bb}} * 15:12 ejegg: payments-wiki upgraded from {{Gerrit|1ad85e6c}} to {{Gerrit|064a770e}} * 01:19 ejegg: donorwiki upgraded from {{Gerrit|1ad85e6c}} to {{Gerrit|064a770e}} * 00:26 dwisehaupt: cloning new frdb frdb1008 from frdb2005 == 2026-04-07 == * 18:34 wfan: civicrm upgraded from {{Gerrit|9104e70b}} to {{Gerrit|6f762e29}} == 2026-04-06 == * 20:42 ejegg: re-enabled recurring donation charge job * 20:33 wfan: donorwiki upgraded from {{Gerrit|c2d03117}} to {{Gerrit|1ad85e6c}} * 20:32 wfan: payments-wiki upgraded from {{Gerrit|80cda166}} to {{Gerrit|1ad85e6c}} * 16:53 ejegg: disabled recurring donations charge job while diagnosing gr4vy routing errors * 16:03 ejegg: civicrm upgraded from {{Gerrit|4ee11209}} to {{Gerrit|9104e70b}} == 2026-04-03 == * 00:04 wfan: civicrm upgraded from {{Gerrit|49f541cd}} to {{Gerrit|4ee11209}} == 2026-04-02 == * 21:38 cstone: payments-wiki upgraded from {{Gerrit|86bec442}} to {{Gerrit|80cda166}} * 05:24 eileen: civicrm upgraded from {{Gerrit|c512abc6}} to {{Gerrit|49f541cd}} * 02:39 eileen: civicrm upgraded from {{Gerrit|bbed1291}} to {{Gerrit|c512abc6}} * 02:16 eileen: SmashPig upgraded from {{Gerrit|9af71a7c}} to {{Gerrit|18ea746a}} == 2026-04-01 == * 18:57 eileen: civicrm upgraded from {{Gerrit|a1bf4768}} to {{Gerrit|bbed1291}} * 04:11 eileen: civicrm upgraded from {{Gerrit|11a2f9ab}} to {{Gerrit|a1bf4768}} * 03:18 ejegg: payments-wiki upgraded from {{Gerrit|02bf54b0}} to {{Gerrit|86bec442}} == 2026-03-31 == * 22:03 jgleeson: tools upgraded from {{Gerrit|9985e723}} to {{Gerrit|986f7f83}} * 20:16 eileen: civicrm upgraded from {{Gerrit|c3cc3562}} to {{Gerrit|b468301c}} * 18:40 jgleeson: tools upgraded from {{Gerrit|161049ac}} to {{Gerrit|9985e723}} * 17:39 ejegg: Standalone (IPN listener) SmashPig upgraded from {{Gerrit|abf8682a}} to {{Gerrit|9af71a7c}} * 16:38 jgleeson: tools upgraded from {{Gerrit|f605b570}} to {{Gerrit|161049ac}} * 16:24 jgleeson: donorwiki updated from {{Gerrit|d79a98b5}} to {{Gerrit|c2d03117}} * 02:46 eileen: civicrm upgraded from {{Gerrit|591bef29}} to {{Gerrit|c3cc3562}} * 01:00 eileen: civicrm upgraded from {{Gerrit|cf871dd3}} to {{Gerrit|591bef29}} == 2026-03-30 == * 23:19 eileen: civicrm upgraded from {{Gerrit|7d299b48}} to {{Gerrit|cf871dd3}} * 21:09 eileen: civicrm upgraded from {{Gerrit|3724cc2d}} to {{Gerrit|7d299b48}} * 20:40 eileen: civicrm upgraded from {{Gerrit|58426b1e}} to {{Gerrit|3724cc2d}} * 14:51 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|545f0b10}} to {{Gerrit|abf8682a}} == 2026-03-28 == * 02:57 ejegg: payments-wiki upgraded from {{Gerrit|d79a98b5}} to {{Gerrit|b239a6b7}} * 02:27 eileen: civicrm upgraded from {{Gerrit|7138d524}} to {{Gerrit|58426b1e}} == 2026-03-27 == * 22:37 eileen: civicrm upgraded from {{Gerrit|c51e98cc}} to {{Gerrit|7138d524}} * 17:50 dwisehaupt: Corrected value: on frdb1005 running the following in mysql to up the buffer pool to ~412G: set global innodb_buffer_pool_size = 431669379072; * 17:34 dwisehaupt: on frdb1005 running the following in mysql to up the buffer pool to ~412G: set global innodb_buffer_pool_size = 421552128; == 2026-03-26 == * 20:00 jgleeson: donorwiki upgraded from {{Gerrit|48cc2e9d}} to {{Gerrit|d79a98b5}} * 19:59 jgleeson: payments-wiki upgraded from {{Gerrit|b387c6ba}} to {{Gerrit|d79a98b5}} * 19:48 ejegg: fundraising civicrm upgraded from {{Gerrit|88b497b8}} to {{Gerrit|db102c77}} * 19:25 larssandergreen: civicrm upgraded from {{Gerrit|26dab2f0}} to {{Gerrit|88b497b8}} * 16:26 jgleeson: SmashPig upgraded from {{Gerrit|5d8a0330}} to {{Gerrit|545f0b10}} * 02:13 eileen: civicrm upgraded from {{Gerrit|a752f9e8}} to {{Gerrit|26dab2f0}} * 00:05 eileen: civicrm upgraded from {{Gerrit|277fe75e}} to {{Gerrit|a752f9e8}} == 2026-03-25 == * 21:53 eileen: civicrm upgraded from {{Gerrit|97319295}} to {{Gerrit|277fe75e}} * 19:58 eileen: civicrm upgraded from {{Gerrit|d3dfedb4}} to {{Gerrit|97319295}} == 2026-03-24 == * 19:32 wfan: payments-wiki upgraded from {{Gerrit|fce3fab5}} to {{Gerrit|b387c6ba}} * 04:24 eileen: civicrm upgraded from {{Gerrit|0aef661f}} to {{Gerrit|b2ed875e}} * 03:38 eileen: config revision changed from {{Gerrit|ded0c289}} to {{Gerrit|16592428}} schedule stripe download * 03:33 eileen: config revision changed from {{Gerrit|79e052e4}} to {{Gerrit|ded0c289}} temporarily disable adyen audit parse - let's fix those misplaced IDs * 01:55 eileen: civicrm upgraded from {{Gerrit|b2c7f1d0}} to {{Gerrit|0aef661f}} * 00:08 eileen: civicrm upgraded from {{Gerrit|80344f51}} to {{Gerrit|b2c7f1d0}} == 2026-03-23 == * 21:30 eileen: config revision changed from {{Gerrit|8c5587f3}} to {{Gerrit|2dd50e7c}} * 18:55 wfan: civicrm upgraded from {{Gerrit|675455b2}} to {{Gerrit|80344f51}} * 17:27 larssandergreen: tools upgraded from {{Gerrit|e60f63b3}} to {{Gerrit|f605b570}} * 15:59 ejegg: civicrm upgraded from {{Gerrit|a2d4b17c}} to {{Gerrit|675455b2}} * 12:25 jgleeson: payments-wiki upgraded from {{Gerrit|48cc2e9d}} to {{Gerrit|91d9eee9}} == 2026-03-20 == * 00:34 eileen: * civicrm upgraded from {{Gerrit|adc36173}} to {{Gerrit|a2d4b17c}} * 00:31 cstone: payments-wiki upgraded from {{Gerrit|f3420a6f}} to {{Gerrit|48cc2e9d}} * 00:27 eileen: config revision changed from {{Gerrit|a7486f6a}} to {{Gerrit|a1a426f3}} * 00:27 eileen: SmashPig upgraded from {{Gerrit|78a8e70a}} to {{Gerrit|5d8a0330}} == 2026-03-19 == * 14:08 damilare: civiproxy upgraded from {{Gerrit|6625c844}} to {{Gerrit|38ba8348}} == 2026-03-17 == * 18:40 jgleeson: donorwiki upgraded from {{Gerrit|4c09db39}} to {{Gerrit|7d1666f9}} * 06:20 eileen: civicrm upgraded from {{Gerrit|7fe14629}} to {{Gerrit|adc36173}} * 05:06 eileen: civicrm upgraded from {{Gerrit|e622a222}} to {{Gerrit|7fe14629}} * 03:48 eileen: civicrm upgraded from {{Gerrit|5360f9ad}} to {{Gerrit|e622a222}} * 02:13 eileen: civicrm upgraded from {{Gerrit|3283e3ca}} to {{Gerrit|e73c6b50}} == 2026-03-15 == * 23:56 eileen: civicrm upgraded from {{Gerrit|dce257f0}} to {{Gerrit|3283e3ca}} * 19:43 eileen: civicrm upgraded from {{Gerrit|a1279ee4}} to {{Gerrit|dce257f0}} == 2026-03-13 == * ish: payments-wiki upgraded from {{Gerrit|f40a1153}} to {{Gerrit|f3420a6f}} == 2026-03-11 == * 21:20 larssandergreen: civicrm upgraded from {{Gerrit|c2c716ca}} to {{Gerrit|a1279ee4}} * 19:15 eileen: civicrm upgraded from {{Gerrit|81baf495}} to {{Gerrit|c2c716ca}} * 07:18 eileen: config revision changed from {{Gerrit|ed2295ab}} to {{Gerrit|a7486f6a}} * 07:02 eileen: civicrm upgraded from {{Gerrit|f418297f}} to {{Gerrit|81baf495}} * 02:28 eileen: civicrm upgraded from {{Gerrit|14e8200e}} to {{Gerrit|da26f37d}} * 00:06 eileen: civicrm upgraded from {{Gerrit|fbb38eda}} to {{Gerrit|14e8200e}} == 2026-03-10 == * 22:08 eileen: civicrm upgraded from {{Gerrit|ef319ea3}} to {{Gerrit|fbb38eda}} * 19:35 eileen: civicrm upgraded from {{Gerrit|773d9fb9}} to {{Gerrit|ef319ea3}} * 06:06 eileen: config revision changed from {{Gerrit|b9bc2a20}} to {{Gerrit|60ef6709}} == 2026-03-09 == * 17:37 damilare: payments-wiki upgraded from {{Gerrit|5b747b97}} to {{Gerrit|f40a1153}} == 2026-03-06 == * 23:39 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|ed16a2ea}} to {{Gerrit|78a8e70a}} * 21:48 ejegg: fundraising civicrm upgraded from {{Gerrit|8aadcd81}} to {{Gerrit|773d9fb9}} * 21:13 ejegg: civicrm fundraising upgraded from {{Gerrit|a1f32ed6}} to {{Gerrit|8aadcd81}} * 20:35 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|217fc7fc}} to {{Gerrit|ed16a2ea}} * 19:18 ejegg: civicrm upgraded from {{Gerrit|f2633c89}} to {{Gerrit|a1f32ed6}} * 03:08 larssandergreen: civicrm upgraded from {{Gerrit|fbac3ce7}} to {{Gerrit|2bae36fa}} * 02:09 ejegg: payments-wiki upgraded from {{Gerrit|9ae5bf60}} to {{Gerrit|5b747b97}} == 2026-03-05 == * 17:53 ejegg: donorwiki upgraded from {{Gerrit|7329b41d}} to {{Gerrit|4c09db39}} * 05:59 eileen: ivicrm upgraded from {{Gerrit|11e5a5d8}} to {{Gerrit|fbac3ce7}} * 05:08 eileen: * civicrm upgraded from {{Gerrit|8bdce85f}} to {{Gerrit|11e5a5d8}} == 2026-03-04 == * 18:52 jgleeson: tools upgraded from {{Gerrit|a3568ffc}} to {{Gerrit|e60f63b3}} * 03:21 ejegg: payments-wiki upgraded from {{Gerrit|5e4939a3}} to {{Gerrit|9ae5bf60}} * 03:20 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|78960c68}} to {{Gerrit|217fc7fc}} == 2026-03-03 == * 20:42 eileen: civicrm upgraded from {{Gerrit|b610f844}} to {{Gerrit|8bdce85f}} * 17:09 dwisehaupt: latest php8.2 updates installed on civi1002 * 06:00 eileen: * civicrm upgraded from {{Gerrit|f4a70c82}} to {{Gerrit|b610f844}} == 2026-03-02 == * 20:40 eileen: cv upgraded from {{Gerrit|dfeedcbe}} to {{Gerrit|f19e0961}} * 18:41 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|8cd1593b}} to {{Gerrit|78960c68}} * 15:13 ejegg: donorwiki upgraded from {{Gerrit|f5d7179a}} to {{Gerrit|7329b41d}} == 2026-02-27 == * 19:00 ejegg: fundraising civicrm upgraded from {{Gerrit|162bbf7c}} to {{Gerrit|f4a70c82}} * 12:24 jgleeson: payments-wiki upgraded from {{Gerrit|f4fd71ff}} to {{Gerrit|a8a9ce78}} * 11:25 jgleeson: payments-wiki upgraded from {{Gerrit|974af222}} to {{Gerrit|f4fd71ff}} == 2026-02-26 == * 22:36 eileen: civicrm upgraded from {{Gerrit|6995dc03}} to {{Gerrit|162bbf7c}} * 21:52 ejegg: civicrm upgraded from {{Gerrit|2a6001d3}} to {{Gerrit|6995dc03}} * 17:44 larssandergreen: civicrm upgraded from {{Gerrit|3ac93e95}} to {{Gerrit|2a6001d3}} * 07:14 eileen: civicrm upgraded from {{Gerrit|c3157fbf}} to {{Gerrit|3ac93e95}} == 2026-02-25 == * 18:33 ejegg: fundraising civicrm upgraded from {{Gerrit|f881e026}} to {{Gerrit|c3157fbf}} * 03:37 eileen: config revision changed from {{Gerrit|390d6434}} to {{Gerrit|a0228e6c}} turn off trustly audit == 2026-02-24 == * 22:48 ejegg: payments-wiki upgraded from {{Gerrit|e5f73610}} to {{Gerrit|974af222}} * 19:47 ejegg: payments-wiki upgraded from {{Gerrit|f5d7179a}} to {{Gerrit|e5f73610}} * 02:06 eileen: config revision changed from {{Gerrit|71c98072}} to {{Gerrit|390d6434}} reenabled trustly audit == 2026-02-23 == * 23:42 eileen: civicrm upgraded from {{Gerrit|f0710864}} to {{Gerrit|f881e026}} * 22:28 ejegg: fundraising civicrm upgraded from {{Gerrit|9d58ce4a}} to {{Gerrit|f0710864}} * 16:43 jgleeson: payments-wiki upgraded from {{Gerrit|0127f2d8}} to {{Gerrit|f5d7179a}} == 2026-02-21 == * 01:15 dwisehaupt: updating localsettings from {{Gerrit|71c98072}} to {{Gerrit|534fbf34}} and syncing civicrm to push update for large_donation_notifications == 2026-02-20 == * 20:38 ejegg: donorwiki upgraded from {{Gerrit|f7a0ee6b}} to {{Gerrit|f5d7179a}} == 2026-02-19 == * 21:26 ejegg: payments-wiki upgraded from {{Gerrit|f7a0ee6b}} to {{Gerrit|0127f2d8}} * 06:37 eileen: civicrm upgraded from {{Gerrit|ac30e19f}} to {{Gerrit|9d58ce4a}} == 2026-02-18 == * 22:21 dwisehaupt: disabling apache2::mod::dump_io on civicrm role for debugging 500 errors after testing. can be re-enabled by reverting commit {{Gerrit|4b1d94399}} - [[phab:T417310|T417310]] * 20:38 eileen: civicrm upgraded from {{Gerrit|f5020a85}} to {{Gerrit|ac30e19f}} * 19:22 eileen: civicrm upgraded from {{Gerrit|66d2e1dd}} to {{Gerrit|f5020a85}} * 17:48 ejegg: donorwiki upgraded from {{Gerrit|488431ec}} to {{Gerrit|f7a0ee6b}} * 16:51 dwisehaupt: enabling apache2::mod::dump_io on civicrm role for debugging 500 errors - [[phab:T417310|T417310]] * 06:02 eileen: civicrm upgraded from {{Gerrit|4c3cdcde}} to {{Gerrit|66d2e1dd}} * 05:33 eileen: config revision changed from {{Gerrit|605c6946}} to {{Gerrit|368156fa}} * 05:21 eileen: civicrm upgraded from {{Gerrit|caad5ab9}} to {{Gerrit|4c3cdcde}} == 2026-02-17 == * 17:34 larssandergreen: payments-wiki upgraded from {{Gerrit|c506d590}} to {{Gerrit|488431ec}} * 17:33 larssandergreen: donorwiki upgraded from {{Gerrit|93e0d03f}} to {{Gerrit|488431ec}} * 03:02 eileen: civicrm upgraded from {{Gerrit|89782fc6}} to {{Gerrit|caad5ab9}} == 2026-02-16 == * 20:16 eileen: civicrm upgraded from {{Gerrit|2b227403}} to {{Gerrit|89782fc6}} == 2026-02-15 == * 23:55 eileen: civicrm upgraded from {{Gerrit|de8252c7}} to {{Gerrit|2b227403}} * 21:07 eileen: civicrm upgraded from {{Gerrit|038f5bca}} to {{Gerrit|de8252c7}} == 2026-02-13 == * 16:03 ejegg: payments-wiki upgraded from {{Gerrit|5793a405}} to {{Gerrit|c506d590}} * away: SmashPig upgraded from {{Gerrit|fea03fcc}} to {{Gerrit|8cd1593b}} == 2026-02-12 == * 22:49 eileen: civicrm upgraded from {{Gerrit|c6c0d453}} to {{Gerrit|038f5bca}} * 21:45 jgleeson: payments-wiki upgraded from {{Gerrit|9dbf0ece}} to {{Gerrit|5793a405}} * 21:15 jgleeson: payments-wiki upgraded from {{Gerrit|6c1a522f}} to {{Gerrit|9dbf0ece}} * 19:03 larssandergreen: tools upgraded from {{Gerrit|645cf5dc}} to {{Gerrit|a3568ffc}} * 00:34 larssandergreen: civicrm upgraded from {{Gerrit|e13111f3}} to {{Gerrit|c6c0d453}} == 2026-02-11 == * 21:12 eileen: civicrm upgraded from {{Gerrit|6e57071a}} to {{Gerrit|e13111f3}} * 06:27 eileen: civicrm upgraded from {{Gerrit|98c325dd}} to {{Gerrit|6e57071a}} * 04:55 larssandergreen: tools upgraded from {{Gerrit|7462b8bd}} to {{Gerrit|645cf5dc}} * 03:37 eileen: civicrm upgraded from {{Gerrit|953cf9f2}} to {{Gerrit|98c325dd}} == 2026-02-09 == * 15:35 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|937d6e40}} to {{Gerrit|fea03fcc}} == 2026-02-08 == * 23:47 eileen: civicrm upgraded from {{Gerrit|bc3a8036}} to {{Gerrit|953cf9f2}} * 20:46 eileen: civicrm upgraded from {{Gerrit|40189afa}} to {{Gerrit|bc3a8036}} == 2026-02-06 == * 15:41 ejegg: fundraising civicrm upgraded from {{Gerrit|06842fbf}} to {{Gerrit|40189afa}} * 06:24 eileen: civicrm upgraded from {{Gerrit|b63d7146}} to {{Gerrit|06842fbf}} == 2026-02-05 == * 15:25 dwisehaupt: all eqiad hosts powered down and ready for relocation. * 15:03 dwisehaupt: starting poweroff of eqiad hosts * 14:48 dwisehaupt: downtimes scheduled for frack eqiad hosts and cross colo replication in prep for rack expansion - [[phab:T403035|T403035]] * 04:30 eileen: civicrm upgraded from {{Gerrit|000dd548}} to {{Gerrit|b63d7146}} == 2026-02-04 == * 23:56 eileen: civicrm upgraded from {{Gerrit|4c2870c5}} to {{Gerrit|000dd548}} * 23:47 cstone: donorwiki upgraded from {{Gerrit|53bfb05b}} to {{Gerrit|93e0d03f}} * 21:51 ejegg: payments-wiki upgraded from {{Gerrit|a09a4f8f}} to {{Gerrit|93e0d03f}} * 21:46 eileen: civicrm upgraded from {{Gerrit|dd10342d}} to {{Gerrit|4c2870c5}} * 21:22 eileen: civicrm upgraded from {{Gerrit|8aa9274a}} to {{Gerrit|dd10342d}} * 03:48 eileen: * civicrm upgraded from {{Gerrit|14c7b7e7}} to {{Gerrit|8aa9274a}} == 2026-02-03 == * 21:55 dwisehaupt: as part of dns cleanup, we have removed the old civi1002.wikimedia.org entry. folks should be using civicrm.wm.o but there is a chance of super old bookmarks still being around. we can reinstate if it becomes an issue. * 21:05 eileen: config revision changed from {{Gerrit|23b2d9b6}} to {{Gerrit|45d40cf1}} * 21:01 larssandergreen: civicrm upgraded from {{Gerrit|10ab3659}} to {{Gerrit|14c7b7e7}} * 05:49 eileen: config revision changed from {{Gerrit|f348441d}} to {{Gerrit|23b2d9b6}} * 05:48 eileen: civicrm upgraded from {{Gerrit|a097bb3d}} to {{Gerrit|10ab3659}} * 01:16 wfan: civicrm upgraded from {{Gerrit|5b9f3cd4}} to {{Gerrit|a097bb3d}} * 00:24 wfan: donorwiki upgraded from {{Gerrit|3ffc70f0}} to {{Gerrit|53bfb05b}} == 2026-02-02 == * 23:50 wfan: payments-wiki upgraded from {{Gerrit|c035aa84}} to {{Gerrit|53bfb05b}} * 20:22 cstone: civicrm upgraded from {{Gerrit|f91f955b}} to {{Gerrit|5b9f3cd4}} * 18:55 ejegg: standalone (IPN listener) SmashPig upgraded from {{Gerrit|b42f2de4}} to {{Gerrit|937d6e40}} * 17:01 wfan: payments-wiki upgraded from {{Gerrit|c5cadd72}} to {{Gerrit|c035aa84}} * 15:26 ejegg: fundraising civicrm upgraded from {{Gerrit|611d18de}} to {{Gerrit|f91f955b}} == 2026-01-30 == * 17:02 larssandergreen: civicrm upgraded from {{Gerrit|79e4424e}} to {{Gerrit|611d18de}} * 04:18 cstone: civicrm upgraded from {{Gerrit|fe1af57a}} to {{Gerrit|79e4424e}} * 02:51 eileen: civicrm upgraded from {{Gerrit|ff772bee}} to {{Gerrit|fe1af57a}} * 02:28 eileen: civicrm upgraded from {{Gerrit|bcf976ae}} to {{Gerrit|ff772bee}} == 2026-01-29 == * 22:16 eileen: civicrm upgraded from {{Gerrit|5d121c63}} to {{Gerrit|bcf976ae}} * 20:17 eileen: civicrm upgraded from {{Gerrit|ebcfd009}} to {{Gerrit|5d121c63}} * 20:09 eileen: civicrm upgraded from {{Gerrit|5c065c4e}} to {{Gerrit|ebcfd009}} * 18:08 wfan: payments-wiki upgraded from {{Gerrit|81d9f614}} to {{Gerrit|c5cadd72}} == 2026-01-28 == * 19:36 jgleeson: SmashPig upgraded from {{Gerrit|96a6224d}} to {{Gerrit|b42f2de4}} * 18:53 jgleeson: civicrm upgraded from {{Gerrit|56c222da}} to {{Gerrit|5c065c4e}} * 18:53 jgleeson: SmashPig upgraded from {{Gerrit|96a6224d}} to {{Gerrit|b42f2de4}} * 14:41 jgleeson: payments-wiki upgraded from {{Gerrit|24915bdb}} to {{Gerrit|81d9f614}} * 07:30 eileen: civicrm upgraded from {{Gerrit|600b21a6}} to {{Gerrit|56c222da}} * 06:51 eileen: * civicrm upgraded from {{Gerrit|32f9a10d}} to {{Gerrit|600b21a6}} * 04:35 eileen: config revision changed from {{Gerrit|ed0808a9}} to {{Gerrit|ef6ef5f2}} * 03:42 eileen: civicrm upgraded from {{Gerrit|64267a34}} to {{Gerrit|32f9a10d}} * 02:13 eileen: civicrm upgraded from {{Gerrit|7299615a}} to {{Gerrit|64267a34}} == 2026-01-27 == * 21:12 larssandergreen: tools upgraded from {{Gerrit|84323460}} to {{Gerrit|7462b8bd}} * 05:57 cstone: civicrm upgraded from {{Gerrit|19f94835}} to {{Gerrit|75f443b5}} == 2026-01-26 == * 17:15 damilare: smashpig upgraded from {{Gerrit|8b4ebf34}} to {{Gerrit|96a6224d}} * 16:32 larssandergreen: tools upgraded from {{Gerrit|c75f7625}} to {{Gerrit|84323460}} * 01:43 eileen: config revision changed from {{Gerrit|2f71107f}} to {{Gerrit|ed0808a9}} switch to php dlocal downloader (now weekend is mostly over) - == 2026-01-25 == * 21:47 eileen: config revision changed from {{Gerrit|23023984}} to {{Gerrit|2f71107f}} == 2026-01-24 == * 02:19 cstone: civicrm upgraded from {{Gerrit|f7064a46}} to {{Gerrit|19f94835}} * 00:55 bd808: Testing #wikimedia-fundraising SAL integration ([[phab:T415389|T415389]]) <noinclude>[[Category:SAL]]</noinclude> avorthnj5o5tz1kdwks9gdvtzog45xh Nova Resource:Tools.wikimonitor/SAL 498 460126 2445252 2438701 2026-08-09T14:25:02Z Stashbot 7414 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment | ref=v1.6.2 2445252 wikitext text/x-wiki === 2026-08-09 === * 14:25 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.6.2 === 2026-07-19 === * 16:04 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/67) has been submitted by praffq-dev has been merged ([[phab:T425925|T425925]]) * 14:11 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/67) has been submitted by praffq-dev on GitHub ([[phab:T425925|T425925]]) === 2026-07-17 === * 23:32 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Deployment completed successfully {{!}} ref=v1.6.1 * 23:32 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.6.1 * 23:32 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.6.1 * 23:30 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.6.1 === 2026-07-02 === * 15:29 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Deployment completed successfully {{!}} ref=v1.5.7 (for fix v1.5.6) * 15:29 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.7 (for fix v1.5.6) * 15:29 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.7 (for fix v1.5.6) * 15:28 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.7 (for fix v1.5.6) === 2026-06-29 === * 16:06 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Deployment completed successfully {{!}} ref=v1.5.6 * 16:06 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.6 * 16:06 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.6 * 16:04 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.6 === 2026-05-26 === * 15:12 wmbot~gergesshamon@tools-bastion-15: Restarting buildservice === 2026-05-10 === * 13:22 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/57) has been submitted by Gerges has been merged ([[phab:T424152|T424152]]) * 13:21 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/57) has been submitted by Gerges on GitHub ([[phab:T424152|T424152]]) === 2026-05-07 === * 16:23 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.5 * 16:23 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.5 * 16:21 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.5 === 2026-05-06 === * 18:42 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/53) has been submitted by praffq-dev has been merged ([[phab:T422194|T422194]]) * 18:42 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/56) has been submitted by �[200~praffq-dev has been merged ([[phab:T422194|T422194]]) * 17:23 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/56) has been submitted by Gerges has been merged ([[phab:T425534|T425534]]) * 17:12 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/56) has been submitted by Gerges on GitHub ([[phab:T425534|T425534]]) === 2026-05-03 === * 09:31 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.4 * 09:31 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.4 * 09:30 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.4 === 2026-05-01 === * 16:35 wmbot~gergesshamon@tools-bastion-15: setup <noinclude>[[Category:SAL]]</noinclude> cxzj1swoq0mnuypi9j9vpavutcjpoo5 2445253 2445252 2026-08-09T14:26:18Z Stashbot 7414 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully | ref=v1.6.2 2445253 wikitext text/x-wiki === 2026-08-09 === * 14:26 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.6.2 * 14:25 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.6.2 === 2026-07-19 === * 16:04 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/67) has been submitted by praffq-dev has been merged ([[phab:T425925|T425925]]) * 14:11 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/67) has been submitted by praffq-dev on GitHub ([[phab:T425925|T425925]]) === 2026-07-17 === * 23:32 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Deployment completed successfully {{!}} ref=v1.6.1 * 23:32 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.6.1 * 23:32 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.6.1 * 23:30 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.6.1 === 2026-07-02 === * 15:29 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Deployment completed successfully {{!}} ref=v1.5.7 (for fix v1.5.6) * 15:29 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.7 (for fix v1.5.6) * 15:29 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.7 (for fix v1.5.6) * 15:28 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.7 (for fix v1.5.6) === 2026-06-29 === * 16:06 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Deployment completed successfully {{!}} ref=v1.5.6 * 16:06 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.6 * 16:06 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.6 * 16:04 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.6 === 2026-05-26 === * 15:12 wmbot~gergesshamon@tools-bastion-15: Restarting buildservice === 2026-05-10 === * 13:22 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/57) has been submitted by Gerges has been merged ([[phab:T424152|T424152]]) * 13:21 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/57) has been submitted by Gerges on GitHub ([[phab:T424152|T424152]]) === 2026-05-07 === * 16:23 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.5 * 16:23 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.5 * 16:21 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.5 === 2026-05-06 === * 18:42 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/53) has been submitted by praffq-dev has been merged ([[phab:T422194|T422194]]) * 18:42 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/56) has been submitted by �[200~praffq-dev has been merged ([[phab:T422194|T422194]]) * 17:23 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/56) has been submitted by Gerges has been merged ([[phab:T425534|T425534]]) * 17:12 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/56) has been submitted by Gerges on GitHub ([[phab:T425534|T425534]]) === 2026-05-03 === * 09:31 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.4 * 09:31 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.4 * 09:30 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.4 === 2026-05-01 === * 16:35 wmbot~gergesshamon@tools-bastion-15: setup <noinclude>[[Category:SAL]]</noinclude> 7vz50b67fa03lfi7chqi4k5qrhfdtlr 2445254 2445253 2026-08-09T14:26:18Z Stashbot 7414 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice | ref=v1.6.2 2445254 wikitext text/x-wiki === 2026-08-09 === * 14:26 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.6.2 * 14:26 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.6.2 * 14:25 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.6.2 === 2026-07-19 === * 16:04 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/67) has been submitted by praffq-dev has been merged ([[phab:T425925|T425925]]) * 14:11 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/67) has been submitted by praffq-dev on GitHub ([[phab:T425925|T425925]]) === 2026-07-17 === * 23:32 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Deployment completed successfully {{!}} ref=v1.6.1 * 23:32 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.6.1 * 23:32 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.6.1 * 23:30 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.6.1 === 2026-07-02 === * 15:29 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Deployment completed successfully {{!}} ref=v1.5.7 (for fix v1.5.6) * 15:29 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.7 (for fix v1.5.6) * 15:29 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.7 (for fix v1.5.6) * 15:28 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.7 (for fix v1.5.6) === 2026-06-29 === * 16:06 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Deployment completed successfully {{!}} ref=v1.5.6 * 16:06 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.6 * 16:06 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.6 * 16:04 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.6 === 2026-05-26 === * 15:12 wmbot~gergesshamon@tools-bastion-15: Restarting buildservice === 2026-05-10 === * 13:22 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/57) has been submitted by Gerges has been merged ([[phab:T424152|T424152]]) * 13:21 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/57) has been submitted by Gerges on GitHub ([[phab:T424152|T424152]]) === 2026-05-07 === * 16:23 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.5 * 16:23 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.5 * 16:21 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.5 === 2026-05-06 === * 18:42 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/53) has been submitted by praffq-dev has been merged ([[phab:T422194|T422194]]) * 18:42 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/56) has been submitted by �[200~praffq-dev has been merged ([[phab:T422194|T422194]]) * 17:23 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/56) has been submitted by Gerges has been merged ([[phab:T425534|T425534]]) * 17:12 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/56) has been submitted by Gerges on GitHub ([[phab:T425534|T425534]]) === 2026-05-03 === * 09:31 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.4 * 09:31 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.4 * 09:30 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.4 === 2026-05-01 === * 16:35 wmbot~gergesshamon@tools-bastion-15: setup <noinclude>[[Category:SAL]]</noinclude> fev99g9jr7bfx9tspau5cjgkrcclmf9 2445255 2445254 2026-08-09T14:26:25Z Stashbot 7414 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Deployment completed successfully | ref=v1.6.2 2445255 wikitext text/x-wiki === 2026-08-09 === * 14:26 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Deployment completed successfully {{!}} ref=v1.6.2 * 14:26 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.6.2 * 14:26 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.6.2 * 14:25 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.6.2 === 2026-07-19 === * 16:04 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/67) has been submitted by praffq-dev has been merged ([[phab:T425925|T425925]]) * 14:11 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/67) has been submitted by praffq-dev on GitHub ([[phab:T425925|T425925]]) === 2026-07-17 === * 23:32 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Deployment completed successfully {{!}} ref=v1.6.1 * 23:32 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.6.1 * 23:32 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.6.1 * 23:30 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.6.1 === 2026-07-02 === * 15:29 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Deployment completed successfully {{!}} ref=v1.5.7 (for fix v1.5.6) * 15:29 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.7 (for fix v1.5.6) * 15:29 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.7 (for fix v1.5.6) * 15:28 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.7 (for fix v1.5.6) === 2026-06-29 === * 16:06 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Deployment completed successfully {{!}} ref=v1.5.6 * 16:06 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.6 * 16:06 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.6 * 16:04 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.6 === 2026-05-26 === * 15:12 wmbot~gergesshamon@tools-bastion-15: Restarting buildservice === 2026-05-10 === * 13:22 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/57) has been submitted by Gerges has been merged ([[phab:T424152|T424152]]) * 13:21 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/57) has been submitted by Gerges on GitHub ([[phab:T424152|T424152]]) === 2026-05-07 === * 16:23 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.5 * 16:23 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.5 * 16:21 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.5 === 2026-05-06 === * 18:42 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/53) has been submitted by praffq-dev has been merged ([[phab:T422194|T422194]]) * 18:42 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/56) has been submitted by �[200~praffq-dev has been merged ([[phab:T422194|T422194]]) * 17:23 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/56) has been submitted by Gerges has been merged ([[phab:T425534|T425534]]) * 17:12 wmbot~gergesshamon@tools-bastion-15: A pull request (https://github.com/wiki-connect/wikimonitor/pull/56) has been submitted by Gerges on GitHub ([[phab:T425534|T425534]]) === 2026-05-03 === * 09:31 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Restarting buildservice {{!}} ref=v1.5.4 * 09:31 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Build triggered successfully {{!}} ref=v1.5.4 * 09:30 wmbot~gergesshamon@tools-bastion-15: [DEPLOY] Starting deployment {{!}} ref=v1.5.4 === 2026-05-01 === * 16:35 wmbot~gergesshamon@tools-bastion-15: setup <noinclude>[[Category:SAL]]</noinclude> t6m18fdgowv335ktkbdqg0jw15w54te Wikidata Query Service/Migration/Development With Local Wiki 0 460538 2445270 2445164 2026-08-10T07:22:42Z Arthur Taylor (WMDE) 37356 spelling mistake 2445270 wikitext text/x-wiki == Testing WikibaseQualityConstraints with a local Wikibase and local QLever == In order to test changes to WikibaseQualityConstraints (and any other features that require a development wiki), it is useful to be able to end-to-end test features that make Sparql queries. One approach for such testing is to run the QLever server locally, and point the locally wiki at a locally-running wdqs-proxy server instance. === Create a new local wiki === Assuming you are using mwcli for local development, you can create a new wiki for your QueryService testing. Will we use the wiki name <code>qswikidev</code> - you will need to ensure that your LocalSettings.php loads Wikibase, QualityConstraints and any other required extensions for the <code>qswikidev</code> database before running <code>update</code> (the last line of the following block): $ mw docker mediawiki install --dbtype=mysql --dbname=qswikidev $ MW_DB=wikidatawikidev mw docker mediawiki exec -- php maintenance/run.php AddSite --wiki qswikidev qswikidev mylocalwikis --interwiki-id qswikidev --navigation-id qswikidev --pagepath '<nowiki>http://qswikidev.mediawiki.mwdd.localhost:8080/w/index.php?title=$1'</nowiki> --filepath '<nowiki>http://qswikidev.mediawiki.mwdd.localhost:8080/w/$1'</nowiki> --language de --server '<nowiki>http://qswikidev.mediawiki.mwdd.localhost:8080'</nowiki> $ mw docker mediawiki mwscript MW_DB=qswikidev -- update --quick If for whatever reason you need to drop the wiki and start over, you can do that with: $ mw docker mediawiki exec -- php maintenance/run.php MwSql --status --wiki qswikidev --query "DROP DATABASE qswikidev;" === Import Constraints and properties / items from Wikidata.org === To make the QualityConstraints feature work, we need appropriate properties and items for the constraints, and the correct configuration for our local wiki. We can import the data from production with the following maintenance scripts: $ mw docker mediawiki exec -- MW_DB=qswikidev php maintenance/run.php ./extensions/WikibaseQualityConstraints/maintenance/ImportConstraintEntities.php Save the output of the script into a file `WDQCPropertySettings_qswikidev.php` in your local Wiki folder, and make sure that is included in DefaultLocalSettings.php for the qswikidev wiki. $ mw docker mediawiki exec -- MW_DB=qswikidev php maintenance/run.php ./extensions/WikibaseQualityConstraints/maintenance/ImportConstraintStatements.php === Running QLever locally === Details of how the QLever service can be run locally are detailed at [[Wikidata Query Service/Migration/Development Infrastructure]] . Per the README at <nowiki>https://gitlab.wikimedia.org/repos/wikidata-platform/triplestores</nowiki> $ sh <(curl -L <nowiki>https://nixos.org/nix/install</nowiki>) --no-daemon $ . .nix-profile/etc/profile.d/nix.sh setup the nix config: # .config/nix/nix.conf experimental-features = nix-command flakes Download and build the QLever server $ nix flake new qlever -t git+<nowiki>https://gitlab.wikimedia.org/repos/wikidata-platform/triplestores#qlever</nowiki> $ cd qlever $ nix build . Dump the local wiki data as <code>.nt</code>: $ mw docker mediawiki exec -- MW_DB=qswikidev php maintenance/run.php ./extensions/Wikibase/repo/maintenance/dumpRdf.php --format=nt > /tmp/dump.nt and import into the server, after deleting header and trailer mess (the maintenance script includes some debug info in stdout): cat /tmp/dump.nt | ../result/bin/IndexBuilderMain -m 1G -F ttl -f - -i wikidata -s ../conf/wikidata.settings.json and run the server: ../result/bin/ServerMain --index-basename wikidata --port 7001 --memory-max-size 32G --cache-max-size 16G --default-query-timeout 2000s -a localhost ==== wdqs-proxy server ==== We also need to setup the wdqs-proxy server to reproduce the setup that we see in production: $ git clone <nowiki>https://gitlab.wikimedia.org/repos/wikidata-platform/wdqs/wdqs-proxy</nowiki> $ cd wdqs-proxy Create an env file: $ cat > envfile << EOF QUARKUS_LOG_LEVEL=INFO QUARKUS_LOG_CONSOLE_JSON=true QUARKUS_LOG_CONSOLE_JSON_LOG_FORMAT=ECS QUARKUS_HTTP_PORT=8090 QUARKUS_REST_CLIENT_DOWNSTREAM_SPARQL_ENDPOINT_URL=<nowiki>http://localhost:7001</nowiki> QUARKUS_REST_CLIENT_EVENTGATE_ENDPOINT_URL=<nowiki>http://eventgate-analytics.discovery.wmnet:4592</nowiki> QUERY_REWRITER_MW_SPARQL_ENABLED=true QUERY_REWRITER_MW_SPARQL_ENDPOINT=<nowiki>https://query.wikidata.org/mwsparql</nowiki> QUERY_REWRITER_FEDERATION_ENABLED=true QUERY_REWRITER_FEDERATION_ALLOWLIST_PATH=/srv/app/config/allowlist.json QUERY_REWRITER_PATH_SEARCH_ENABLED=true QUERY_REWRITER_PATH_SEARCH_MAX_DEPTH=5 EVENTGATE_STREAM=wdqs_internal.sparql_query.v2 EVENTGATE_GRAPH_NAME=wikidata_main WDQS_PROXY_METRICS_PREFIX=wdqs.proxy. EOF Then run the proxy server: $ sudo docker run  --network=host --env-file envfile -v ./src/test/resources/allowlist.json:/srv/app/config/allowlist.json docker-registry.wikimedia.org/repos/wikidata-platform/wdqs/wdqs-proxy:v0.4.0 === Confirm everything is running === You can check the server is running with the following query: $ curl -X POST <nowiki>http://localhost:8090/sparql?access-token=localhost</nowiki> -H "Content-Type: application/sparql-query" --data-binary "SELECT (COUNT(*)  AS ?count) WHERE {?s ?p ?o}" eujnqv9fghjpvausuhckqgmtotvmngr